Why Can’t Most Labs Use Their Own Data for Machine Learning?
Because the workflow lets the most valuable data slip through uncaptured. Every sample movement, reagent scan, operator decision, and instrument cycle generates exactly the information machine learning needs. But the workflow wasn’t designed to capture it as structured, linked, computable data. It captures what audits require and discards the rest.
Why Does the Most Useful Data Never Get Captured?
Because lab workflows were designed to record outcomes, not to capture the operational events that produce them. The LIMS logs that a sample made it through a protocol, and perhaps there is reagent tracking as well. It doesn’t log the twelve workflow events between accessioning and result that determine whether the data is trustworthy, timely, or reproducible.
That gap exists at every layer. Instrument software captures run parameters, but in proprietary formats locked inside the instrument’s own reporting. Reagent usage gets documented in logbooks after the run, disconnected from the sample, the operator, and the timestamp it belongs with. Operator decisions happen at the bench and go unrecorded entirely, or land in free-text fields that are open to interpretation.
The information AI needs (sample history, experimental goals, reagent lot, instrument settings, environmental conditions, workarounds, what happened and in what sequence), all of that exists during execution. The workflow just doesn’t capture it. It passes through the operation and disappears.
This is not a data quality problem. The data isn’t dirty. It was never collected.
The Fix: Making the Workflow the Data Source
The data AI needs already exists inside the lab’s operations and in the mind of your technicians. The fix captures it there. Instead of layering analytics on top of documentation-grade records, the workflow itself becomes the primary data source: structured, linked, machine-readable data produced as a direct output of doing the work. No separate cleaning step. No retrospective assembly. The workflow runs, and the data is ready.
Blanchard Strategic’s Approach:
The work starts by identifying the gap between the data required to reach your AI goals, and the data the current workflow actually produces. That gap exposes where the workflow needs new capture points, where existing data needs structure, and where system boundaries prevent information from linking into a complete operational picture.
Comparing the data that machine learning, analytics, or AI use cases need against what the current workflow actually captures, in what format, and with what linkages.
Identifying the points in the operational workflow where structured data capture replaces manual recording, free-text entry, or after-the-fact documentation.
Designing the keys, identifiers, and relationships that connect samples, reagents, operators, instruments, and outcomes into a single queryable structure.
Defining what each system owns, how data moves between LIMS, instruments, and operational platforms, and where handoffs need explicit mapping rather than assumptions.
Testing whether the redesigned data architecture actually supports the intended AI or analytics use cases before the organization invests in building models or platforms.
When Is This Work Useful?
This engagement fits labs where an AI, machine learning, or advanced analytics initiative depends on operational data that the current workflow generates but doesn’t capture in a usable form. Data structured for machine learning is also easier for humans to query, analyze, and act on.
Written by Megan Blanchard, M.S., Principal Systems Architect at Blanchard Strategic Systems.
