Skip to content

We build AI agents you can check

Two weeks, €3,000 fixed, one task — an agent doing that task on your own data, with a number on how often it gets it right and a list of what it gets wrong. If it does not pay for itself, we say so on day three and there is nothing to pay.

Agents we have shipped

A chatbot answers and forgets. An agent is handed a job, fetches what it needs, works the steps, and comes back with a draft, a filled-in record, a finished report — for a person to approve or reject. Getting that to work in a demo is easy. Getting it to behave a month later, unattended, on inputs nobody thought to try, is the job. So we score them on real test sets before they ship, keep an audit trail of what they did, and pressure-test them against the ways they get tricked.

Multi-agent

A six-agent pipeline

OpenClaw runs insurance claims through six agents with a Run → Eval → Improve loop, and an LLM judge scores every pass.

View the code ↗
Auditable

An on-chain agent you can check

Trustodian enforces a spending mandate and leaves a trail you can audit after the fact. Second place at The Money Agent Hackathon.

Adversarial-tested

Built to resist attacks

Our BitGN agent has to finish ordinary tasks while refusing the prompt-injection traps hidden inside them, and it's scored by what it actually changed.

View the code ↗
In production

Running for real users

A nutrition agent lives in a live chat with memory and context, shipped only after we widened its eval set. A team of agents reads location data to write sourced expansion briefs for a retail-analytics product. And an offer agent builds a uniform manufacturer's client offers from their own catalogue, showing the plan before it spends anything.

What goes into an agent that holds up

Three pieces, in the order you meet them: build one agent for one task and prove it pays, give it access to what the company actually knows, and keep measuring what it gets wrong before a customer finds it. Skip any of the three and you get a demo nobody could turn into a system. Each one is also a service you can buy on its own.

  • 01

    Agent Build

    Two weeks, one task, one working agent on your own data. The first days settle which task pays for itself and end on a go/no-go — if the answer is no, there is nothing to pay. The rest builds the winner against a test set of your real cases. Code, test set and numbers are yours.

    Proven by

    How the build works →
  • 02

    Knowledge Base (RAG)

    Your own documents start answering questions, and every answer links back to where it came from. We agree on a test set up front and show you the quality numbers before handover.

    Proven by

    How we build it →
  • 03

    Evals / QA

    We build test sets from your real cases and score your AI the same way on every release, so a drop in quality shows up before your users run into it.

    Proven by

    How evals work →

Built for one industry

Offer sheets for uniform makers

A one-line brief in chat comes back as a branded offer sheet. It uses the manufacturer's own garments and the client's brand colours, in their sheet format, in Serbian or English.

It has been running on their real client work since July 2026.

See what it does →
An offer sheet for kitchen staff: a pair in the uniform, colour palette, detail close-ups and materials

Built by researchers who ship

PhD-trained researchers and engineers who compete in the open and publish what they build. Each line below is a result someone other than us can look at.

Not sure an agent is the right answer?

Sometimes it is not. If the task runs twice a month, the arithmetic will not work. If nobody inside your company will own the agent after we leave, it will be switched off by winter. If what you want is a chat widget on your site, buy one — it is a solved product at software prices. The first days of a build are a go/no-go, and no is a real answer: if none of your candidate tasks pays for itself, we say so on day three and there is nothing to pay.