Skip to content
← All projects

Market Analyst for a Chemical Manufacturer

In ProductionMarch 1, 2026Ongoing

Choosing what a chemical plant should make next is a bet with capital equipment behind it, and it begins as desk work: a handful of specialists reading supplier spreadsheets, catalogues and price lists, reconciling them by hand, and writing up which products look worth a production line. It is slow, and the slowness is not the worst part. Months later, when someone asks why a particular product was recommended, the answer is usually a person's memory. This analyst was built so that question always has an answer.

Every figure keeps the row it came from

Each value the pipeline produces carries its own provenance: the source it was read from and the evidence behind it, recorded at the row level rather than at the level of a file or a run. That is what makes a later challenge answerable. A number in a spreadsheet is indistinguishable from a number somebody typed; a number with a source row and a quote attached can be re-checked without trusting anyone's recollection, including ours.

Gates are named, and nothing drops silently

Every check the data passes through is a named gate with a severity of its own. What fails is quarantined or rejected explicitly, and the counts appear in the report for that run. The discipline that matters here is negative: a row is never quietly discarded. Silent drops are how a pipeline comes to look cleaner than the data underneath it, and how a screening result ends up resting on records nobody knows were removed.

Matched by registry number, not by name
Compounds are resolved through their CAS registry numbers against public chemistry databases. The same substance travels under several trade and systematic names, so name matching quietly merges things that are not the same and splits things that are.
A queue where a person decides
Anything a gate cannot settle goes to a manual-resolution queue instead of being guessed at. The edge cases are exactly the cases where a confident wrong answer is most expensive, so they are routed to a human rather than resolved by the model.

Messy input treated as the normal case

The input is exported spreadsheets, which means merged cells, unit columns that hold three different units, headers repeated mid-sheet, and identifiers that are only sometimes filled in. The pipeline ingests that as its ordinary case rather than as an exception: it cleans the sheets, extracts the compounds, and aggregates market and price evidence around them. Nothing about the workflow assumes a clean file, because a clean file has never arrived.

Where the fast build fails an audit

Let the model read the documents and report the number
This is the quickest route to a figure and it fails the first audit for one reason: when someone asks where the number came from, the model said so is not an answer. Provenance is not a feature added afterwards — it either survives every hop from source row to final figure or it does not exist.

Months to hours, with the person kept in

The point of the build was to take primary screening from months of work by several specialists down to repeatable analysis measured in hours. The word doing the work in that sentence is repeatable: the same inputs give the same outputs, and a rerun is cheap enough that a changed assumption can be tested rather than argued about. Human verification was not removed to get there. The gates and the resolution queue exist so that a specialist spends their time on the cases that need judgement instead of on the reconciliation that does not.

What it proves

Recommends what to produce next on evidence a buyer can re-check: every figure traced to its source row and cleared by named quality gates.

Category

Data Analytics

Built with

React · TypeScript · Supabase · Postgres · Edge Functions · Quality Gates

Need something similar?

The cheapest way in is two weeks. The first days work out which task would pay for itself in your processes; the rest builds that agent on your own data and measures it.