Researchers are using large language models to unravel the coded language of 17th century alchemical texts, marking a new intersection between artificial intelligence and historical scholarship. The project applies LLMs to trace the transmission of alchemical knowledge across Europe, decoding encrypted letters that have stumped historians for decades.

What You Need to Know

Alchemical manuscripts often used secret symbols and vague language to protect proprietary knowledge. Traditional manual transcription is painstakingly slow. Using LLMs, researchers can process thousands of pages in minutes, identifying patterns that would escape human eyes. This approach could fundamentally change how historians study premodern science.

The Challenge of Alchemical Secrecy

Alchemists in the 17th century deliberately obscured their methods, using encoded terms and cryptic diagrams to hide recipes for transmutation, medicines and philosophical insights. Decoding these texts requires not only linguistic skill but also understanding of the esoteric context in which they were written. LLMs, trained on vast corpora of historical and scientific literature, can fill gaps by predicting likely meanings for ambiguous symbols.

The research focuses on letters exchanged between prominent alchemists across England, Germany and France. Many of these letters were written in a mix of Latin, vernacular languages and invented ciphers. Early results from the model show that it can reconstruct missing words and suggest translations that align with known alchemical theories.

  • Deciphering encrypted text: LLMs identify recurring cipher patterns and map them to probable plaintext.
  • Linking fragmented manuscripts: The model connects related passages across different documents to reconstruct lost workflows.
  • Mapping knowledge networks: By analyzing correspondence, the system shows how alchemical ideas traveled between individuals and regions.

A New Tool for the Humanities

The use of LLMs in historical research goes beyond alchemy. Scholars in digital humanities have applied similar models to analyze medieval poetry, legal documents and early scientific papers. The key advantage is scale: what would take a human expert years, the model can scan in hours. The approach also reduces the risk of confirmation bias, as the model evaluates all data without preconceptions about what the text should say.

Using LLMs for this purpose, however, raises questions about accuracy. The models can hallucinate historical facts or invent plausible-sounding but false translations. Researchers must validate outputs through cross-referencing with known historical records and expert review.

Why This Matters

For historians, this work opens a window into a previously inaccessible layer of scientific history. Alchemical knowledge was a precursor to modern chemistry, and understanding how it spread helps explain the scientific revolution. For the field of AI, the project demonstrates a practical, high-stakes application of language models beyond chatbots and content generation. It shows that LLMs can become partners in scholarly inquiry, provided their limitations are managed.

The broader implication is cultural. By bringing AI into the archive, we may recover knowledge that was intentionally hidden for centuries, reshaping how we view the foundations of science.

Correction: An earlier version of this article misstated the time frame of the manuscripts. They date from the mid-17th century, not the late 1600s.