Founding projects in chemical reaction data, polymer data extraction, infectious disease, and data digitization will bring Notre Dame researchers together to make disciplinary information usable by AI.

The University of Notre Dame's Scientific Artificial Intelligence (SAI) Initiative has launched a set of community projects focused on the central obstacle to using AI in specialized research: most disciplinary knowledge was never organized for machines to read.
Academic publications are formatted for human reading rather than machine ingestion. Raw measurements may be difficult to reconstruct from text alone or placed in unstructured supporting files. Similarly, disciplinary representations and reporting conventions can also hide essential context from a machine such as the conditions of an experiment or standard conventions for plotting results, even when the meaning is clear to an expert reader. The result is that critical information for applying AI to pressing research problems remains effectively out of bounds for contemporary systems.
Groups across Notre Dame are already solving parts of this problem in their own domains. The community projects will connect that work so that each group does not independently rebuild the same extraction machinery. A parsing workflow developed for one literature may contain components that transfer to another, while domain experts determine what information matters and what evidence is needed to interpret it.
One founding project illustrates this combination. Brett Savoie, director of the SAI Initiative and the Coyle Mission Collegiate Professor of Engineering, is working with students Yun-Ke Liu and Bryan Piguave to extract chemical reaction data from the scientific literature.
Reaction data is a bottleneck for applications in drug discovery, materials synthesis, and advanced chemical manufacturing. Determining how to make a molecule of interest remains a fundamental barrier in chemistry, yet relatively little of the field's accumulated knowledge of reactions are available in a form that can be used directly by AI systems.
Chemical structures are commonly embedded as graphics that machines cannot natively interpret. Reaction definitions, conditions, limitations, and characterization may be scattered across figures, tables, experimental sections, and supporting information. Existing datasets reflect these constraints. Some are generated heuristically from reaction templates, while others contain structures without enough context to interpret a reported transformation.
The Savoie Research Group is developing a systematic alternative whose goal is to curate data for every standard chemical reaction class reported in the literature. Liu and Piguave are building specialized workflows based on vision-language models to locate reaction information across the textual and visual formats used in publications and assemble each reaction with its context into consistent records. Unlocking this reaction information could be transformational for ongoing drug development efforts where making molecules of interest is still often the bottleneck.
The data extraction community projects are expected to continue through the summer and fall semester. Participation is open to all experience levels, with several core projects led by teams with experience in this area and new projects being launched. Notre Dame faculty, students, postdoctoral researchers, and staff who are developing extraction workflows, working with difficult domain data, or interested in contributing to a founding project are invited to participate. To join, contact Laura Kresnak at lkresnak@nd.edu.
Originally published by at sai.nd.edu on July 14, 2026.