Insights

From scanned batch records to process intelligence: how BioRaptor uses NVIDIA Nemotron for biologics

September 14, 2026
10 min read
Yaron David, CTO and Co-Founder of BioRaptor
Yaron David
David
CTO and Co-founder
BioRaptor Data Intake screen showing a scanned batch production record with handwritten values mapped to harmonized bioprocess fields and confidence scores

Bringing a new drug to patients is a long, complex and costly process. It starts with discovering a molecule which can have an effect on a target, making sure it works in animal studies, that it is safe and also efficacious. But that is only part of the challenge. A drug must also be manufactured reliably and at scale. For many small-molecule, chemistry-based drugs, manufacturing is relatively straightforward and reproducible. Biologics are different. These can be either an antibody or a short strand of amino acids. This class of drugs is actually produced using living cells, and even small variations in raw materials, equipment, or process conditions can affect the final product. Manufacturing is therefore not merely a step after drug development - it is critical, a defining part of the drug itself.

This is where process development (PD) and Chemistry, Manufacturing and Controls (CMC) people enter the pipeline. While preclinical and clinical teams evaluate safety and efficacy, bioprocess teams define how the biologic will be produced: the cell line or organism, media and feeds, operating conditions, process controls, scale-up strategy, and analytical methods. Their work must ultimately support technology transfer, clinical manufacturing, validation, and commercial production.

During this journey, scientists run hundreds of experiments. Every experiment generates more knowledge - what were the inputs, how the cells behaved and what was the outcome.

For biologics, this process history is critical. Differences in how a batch was prepared and executed can affect productivity, product quality, process robustness, and the ability to reproduce the process at another scale or site.

Much of the executed process is still documented on paper

Paper-based documentation remains prevalent in biopharmaceutical manufacturing. Operators record equipment checks, material lots, additions, samples, setpoint changes, observations, and deviations on forms and executed batch records.

These records are commonly scanned into PDF files for archival, review, and quality workflows. The scan preserves the official record, but it does not make its contents searchable, comparable, or available for analysis.

The same problem exists during process development. Scientists frequently capture execution details in run sheets, spreadsheets, instrument reports, handwritten annotations and filled out SOPs. Formats often differ across scientists, equipment, sites, process stages, and campaigns.

The process history is therefore split across two incomplete views:

  • Bioreactor, analytical, and laboratory systems contain structured measurements.
  • Batch records and run documentation contain the context explaining how the process was actually executed.

Most teams will readily use the first and largely ignore the second.

Process development, accelerated

BioRaptor helps bioprocess teams bring together the complete history of each run: high-frequency sensor data, offline measurements, analytical results, materials, equipment, process events, observations, and outcomes.

Some of the most important context remains trapped inside scanned records:

  • Which material and lot were used
  • When feeds and additions were actually performed
  • Which setpoints were changed - and when
  • What were the P&ID settings
  • Which samples were taken
  • What operators observed
  • Whether an intervention or deviation occurred
  • How the executed process differed from the planned process

Without this context, two batches can appear identical in the structured data even when they were executed differently. Investigations require extensive manual review, comparisons remain incomplete, and machine-learning models learn from only the easiest data to access.

BioRaptor makes the documented execution history part of the process dataset. This allows teams to learn more from every run, resolve problems faster, improve the next experiment, and accelerate progress from process development through manufacturing.

Why conventional OCR is not enough

Extracting this information requires more than converting a scanned page into text.

A conventional OCR system may recognize a number without determining whether it represents a planned setpoint, an actual measurement, a corrected entry, or a material lot. It may read a handwritten note without connecting it to the correct vessel, process stage, field, or timestamp.

Standard named-entity recognition also falls short. Identifying a material name, temperature, or time is not enough - the system must understand how these elements relate to one another and to the batch.

For bioprocess analysis, BioRaptor needs to determine:

  • Which batch, vessel, and process stage each value belongs to
  • Whether a value was planned, observed, or corrected
  • When an addition, sample, alarm, or process change occurred
  • Which material, equipment item, or analytical result is referenced
  • How handwriting, checkboxes, tables, and annotations relate to the surrounding form
  • Whether the record describes the intended process or what was actually executed
  • The exact timing of events

A readable document is not yet usable process data. The information must be structured, placed in context, and connected to the rest of the run.

How BioRaptor uses NVIDIA Nemotron 2 Nano VL

BioRaptor uses NVIDIA Nemotron 2 Nano VL within a document intelligence workflow that converts scanned process documentation into structured process context. We configured a pipeline using Nemotron, without any need for domain specific fine tuning.

Nemotron 2 Nano VL extracts information including:

  • Host, strain, or cell line
  • Media and feed composition
  • Raw materials and lot information
  • Process conditions and setpoints
  • Equipment and vessel information
  • Sampling and addition events
  • Operator observations
  • Interventions and deviations

The extracted information is validated and mapped into BioRaptor’s bioprocess data model. It can then be connected to bioreactor signals, offline measurements, analytical results, annotations, metadata, and process outcomes.

The result is a unified process timeline showing both what the equipment recorded and what the team documented.

Learn from every run: scanned batch records and sensor data pass through NVIDIA Nemotron VL and user validation into BioRaptor for process intelligence

Nemotron 2 Nano VL for bioprocess document interpretation

BioRaptor evaluated Nemotron 2 Nano VL alongside other vision-language models for interpreting real bioprocess documentation.

The clearest advantage was its ability to preserve the relationships expressed through the document’s visual structure.

A handwritten value beside a printed field, a checked box in an equipment section, or a note spanning several table cells can change the interpretation of a run. Other models could often recover the individual words or values but were more likely to lose the relationship between the information and its location on the page.

Nemotron 2 Nano VL provided particular value in three areas:

  • Connecting values to their meaning
    The model was better able to interpret labels, values, tables, checkboxes, and annotations together. This helped distinguish a target from an actual value, an instruction from an executed event, or an original entry from a later correction.
  • Understanding complex and inconsistent layouts
    Process documentation rarely follows one universal format. Records vary across equipment, teams, development stages, sites, and campaigns. Nemotron 2 Nano VL performed well without requiring every document to conform to a rigid template.
  • Producing structured process context
    Rather than returning only a text transcript, Nemotron 2 Nano VL coupled with schema validation tools, supported extraction into a defined bioprocess schema. This reduced the downstream work required to associate extracted information with the correct batch, field, event, and process stage.
  • Analysis speed
    Nemotron 2 Nano VL proved to be x5-x10 faster than other tested off-the shelf models, providing higher throughput and better ROI.

These strengths align with Nemotron 2 Nano VL’s focus on document intelligence, including text and table processing, element parsing, visual reasoning, and complex-layout understanding. NVIDIA also reports leading results on OCRBench v2 for document-oriented tasks such as text recognition, referring, spotting, and parsing. See here for more on NVIDIA Nemotron 2 Nano VL.

From document extraction to process intelligence

Document extraction is only the first step.

Once the information is connected inside BioRaptor, scientists can compare runs without manually reopening hundreds of files. Investigations can include material differences, setpoint changes, sampling events, operator observations, and deviations—not only sensor and assay data.

The enriched dataset can also support better statistical analysis and the building of machine-learning models. Instead of learning from a limited set of numeric variables, models can use a more complete representation of how each process was prepared and executed.

This enables:

  • Faster investigations
  • More complete run-to-run and batch-to-batch comparisons
  • Earlier identification of process differences
  • Better continuity across development, scale-up, and tech transfer
  • More informative machine-learning models
  • Greater value from historical process data

Finding in hours what manual review missed for weeks

“It took me 3 weeks to create the docket the previous time, I’m actually amazed, it could read the handwriting…” - Senior Process Development Engineer at Top-10 manufacturing company

In one customer engagement, the process team had been trying to understand a recurring process issue for several weeks.

BioRaptor collected and connected information that had been distributed across process records, sensor data, analytical measurements, and run metadata. Once the full process context became comparable across runs, the team was able to isolate the root cause within hours.

The value did not come from simply summarizing documents. It came from turning their contents into structured process data and connecting that data to the behavior and outcome of each run.

Building a usable process memory

“I feel like I am standing on the shoulders of giants, our process development team’s data allows us to speed run the optimization process” - Senior Process Development Engineer at a CDMO

Bioprocess organizations already generate substantial process knowledge. The problem is that too much of it remains trapped in documents that cannot be systematically searched, compared, or analyzed.

By combining NVIDIA Nemotron VL with BioRaptor’s bioprocess data model and analytics platform, scanned records become part of the process dataset—not detached evidence reviewed only when something goes wrong.

This gives teams a more complete view of how their processes were actually executed. More importantly, it allows them to learn from every run and use that knowledge to accelerate the next one.

If you would like to see BioRaptor in action, book a demo with us and we’ll show you what’s possible.

Yaron David, CTO and Co-Founder of BioRaptor
Yaron David
David
CTO and Co-founder

Yaron founded BioRaptor out of a life long passion for science and better understanding how things work. Yaron is an MD and holds a PhD in neuroscience and has been developing data intensive platforms throughout his career in both scientific and healthcare settings.

Connect with the author