⏱️ Reading time: 8 min

Novartis cannot feed its artificial intelligence models with decades of chemistry patents and papers: that archive exists in free-form prose, not in a format an algorithm can read. To solve this, CAS, the scientific information division of the American Chemical Society, announced a collaboration with the Swiss pharmaceutical company focused on reaction informatics, that is, on converting that archive into structured data that can be used to train drug-discovery algorithms.

📑 En este artículo
  1. TL;DR
  2. What Is Reaction Informatics
  3. What Happened
  4. Context and History
  5. Technical Details
  6. Impact and Analysis
  7. What’s Next
  8. Frequently Asked Questions
    1. What does Novartis gain from the chemical data curation CAS offers?
    2. How is this different from using SciFinder or Reaxys?
    3. Are other pharmaceutical companies preparing AI-ready reactions like Novartis?
    4. Will CAS release this chemical reaction data to the public?
    5. What is the CAS Registry and why does it matter in this alliance?
  9. References

The agreement targets a well-known bottleneck in computational chemistry: public reaction datasets are small and noisy, while useful knowledge remains locked away in millions of documents that no model can process as-is.

TL;DR

  • CAS and Novartis announced an alliance to structure chemical reaction data suitable for AI.
  • CAS, the American Chemical Society’s division, maintains the CAS Registry and the SciFinder database.
  • Novartis aims to train synthesis-prediction models with reactions that are already normalized and noise-free.
  • Patents and papers rarely present reactions in a format an algorithm can read directly.
  • The agreement adds to the pharmaceutical industry’s race to secure proprietary data against generic AI models.

What Is Reaction Informatics

Reaction informatics is the discipline that converts a chemical reaction described in text (reagents, conditions, catalysts, yield) into a structured record that a machine learning model can process without human intervention, unlike a bibliographic database that only indexes the original document.

The field isn’t new: pharmaceutical AI labs have spent years trying to predict synthesis routes with models like IBM RXN for Chemistry or those trained on the open US patent dataset (USPTO). The limitation was always the same: the most complete version of a reaction tends to be scattered across a paper, a patent, and a lab notebook that no scraper can turn into something trainable without years of manual work by expert chemists.

The same reaction published in two different patents rarely reports its yield in the same format. Foto de Hitesh Choudhary en Unsplash

What Happened

ACS, the parent organization of CAS, confirmed the alliance with Novartis to advance the curation of AI-ready chemical reaction data, though it did not disclose the amounts involved or a delivery timeline. The stated goal is for Novartis to use CAS’s reaction catalog, built from the CAS Registry, as training data for its own drug-discovery models.

The collaboration doesn’t replace the editorial work CAS already does with SciFinder, it extends it. Beyond indexing the document where a reaction appears, each reaction will be tagged with its variables (reagents, solvent, temperature, yield) in a uniform schema that a training pipeline can read without further normalization.

Context and History

CAS was founded as Chemical Abstracts Service in 1907, when the American Chemical Society began summarizing the world’s chemical literature in a single bulletin. Today it manages the CAS Registry, which has assigned a unique identifier to every known chemical substance since 1957, and SciFinder, the search database most industrial and academic chemistry departments use to track reactions, patents, and properties.

A pharmaceutical company like Novartis’s interest in structuring its chemical reaction data isn’t a coincidence. Designing a new drug depends on finding a viable synthesis route: which reagents to combine, in what order, and under what conditions to reach the target molecule without generating toxic byproducts or prohibitive industrial costs. That problem, known as retrosynthesis, is one of the ones that benefits most from automation with models trained on well-labeled reactions, as pharmaceutical AI labs working on the topic have reported for more than a decade.

Technical Details

The mechanism behind reaction informatics isn’t mysterious, but it is costly. A chemistry document describes a reaction in prose: “compound A was dissolved in tetrahydrofuran at 0°C and reagent B was added dropwise.” For a model to use it, someone (a chemist, or an extraction system supervised by one) has to turn that sentence into a table row: the reagent’s structure in SMILES notation, solvent, exact temperature, reaction time, reported yield, and final product, also in SMILES.

That normalization work is the project’s real cost. Databases like Reaxys and SciFinder itself already do a version of this for published literature, but they index the document, not necessarily the reaction in a schema optimized for training: two equivalent reactions described with different terminology can end up as separate records if no one reconciled them. An AI-ready dataset, by contrast, needs deduplication, quality control over reported yield (patents tend to exaggerate it), and a consistent reaction-type taxonomy across the entire corpus.

flowchart TD
A["Patents and chemistry papers"] --> B["Reaction extraction"]
B --> C["Structure normalization (SMILES)"]
C --> D["Curated reaction dataset"]
D --> E["Synthesis prediction model"]

That pipeline explains why a company like CAS, with decades of editorial indexing and staff chemists dedicated to curation, occupies a different position than an AI lab that just scrapes public text: the bottleneck isn’t the model, it’s cleaning the input data.

Impact and Analysis

For Novartis, the agreement fits into a broader pharmaceutical industry strategy: securing preferential access to proprietary data instead of relying on generic models trained on whatever is available on the open web. Chemical reaction data curated by a specialized third party saves Novartis from having to build an in-house team of annotating chemists, a cost that companies of its size typically prefer to outsource to established data providers.

The following table compares the three ways reaction data is currently obtained to train AI models, and why the curated route is gaining ground despite its cost:

MethodData SourceAdvantageLimitation
Free-text miningUnstructured papers and patentsBroad coverage, low initial costHigh noise, missing exact conditions
Commercial databases (Reaxys, SciFinder)Editorially indexed literatureData already normalized per documentPaid access, coverage limited to what’s indexed
Curated chemical reaction data (CAS-Novartis)CAS-Novartis collaborationStructure designed for model trainingInitial scope limited to what both parties agreed on

💭 Key takeaway: the bottleneck in chemical AI was never the model, it was getting enough clean reactions to train it.

The risk for the rest of the industry is concentration: if CAS negotiates similar deals one by one with each major pharmaceutical company, access to quality data for training drug-discovery AI could end up reserved for whoever can pay for it, widening the gap with academic groups or small biotechs that depend on smaller, noisier public datasets.

What’s Next

Neither CAS nor Novartis specified whether the agreement is exclusive or whether CAS will offer the same reaction informatics service to other pharmaceutical companies. The sector’s trend suggests this won’t be the only case: Merck, Pfizer, and other major labs have announced their own investments in drug-discovery AI in recent years, though almost always without detailing where their training data comes from, which is precisely the point this agreement puts on display.

What is predictable is that CAS will use this agreement as a showcase: turning chemical literature into AI-ready data is a service that can be replicated for any industry that depends on reactions, such as agrochemicals, materials, or batteries, not just pharmaceuticals.

If you want to see how structured that data is today, go to cas.org and look up any substance by its CAS Registry number: that’s where you notice the difference between an indexed document and a reaction that’s truly ready to train a model.

📬 Get new articles by email

We only email about big articles (1-2 a month).

Frequently Asked Questions

What does Novartis gain from the chemical data curation CAS offers?

Access to chemical reaction data that’s normalized and verified by specialists, ready to train or fine-tune synthesis-prediction models without building that curation process in-house.

How is this different from using SciFinder or Reaxys?

SciFinder and Reaxys index documents and reactions for a chemist to search manually. This agreement’s chemical data curation structures each reaction as a record with fixed variables, designed for a model to process without manual intervention.

Are other pharmaceutical companies preparing AI-ready reactions like Novartis?

Broadly speaking, yes: labs like Merck and Pfizer invest in artificial intelligence for drug discovery, though they rarely disclose publicly the origin and curation of the reaction data they use to train their models.

Will CAS release this chemical reaction data to the public?

That hasn’t been announced. The agreement is a commercial collaboration between CAS and Novartis, and there’s no indication the curated catalog will be published outside that arrangement.

What is the CAS Registry and why does it matter in this alliance?

The CAS Registry is the system that has assigned a unique identifier to every known chemical substance since 1957. It’s the foundation on which CAS builds its reaction catalog, including the one it’s now preparing for Novartis.

References

  • CAS: scientific information division of the American Chemical Society, responsible for the CAS Registry and the SciFinder database.
  • Novartis: Swiss pharmaceutical company that invests in artificial intelligence applied to drug discovery.
  • American Chemical Society (ACS): parent organization of CAS, a leading authority in chemical publishing and standardization.

📱 Like this content? Follow @programacion on Telegram for daily tech content in Spanish: quick summaries, fresh content every day.

Featured image: Foto de Egor Myznik en Unsplash

Did it work for you? Got a different error? Say so below: questions get answered and help the next reader.

Leave a comment

Andrés Morales

Developer and AI researcher. Writes about language models, frameworks, developer tooling, and open source releases. Covers ML papers, the tech startup ecosystem, and programming trends.

0 Comments

Leave a Reply

Avatar placeholder

Your email address will not be published. Required fields are marked *

You can include code inside <code>…</code> or, for several lines, <pre><code>…</code></pre>.

This site uses Akismet to reduce spam. Learn how your comment data is processed.