Priced by the hour.
Measured by the defect rate.

A dedicated evaluation pod, quoted per productive hour like any other team. The difference is that every queue ships with a published error rate and a gold set you keep.

citation audit · 30-day rolling
actual 0.42%target 0.60%
  • 1,284 citations sampled
  • 94.1% reviewer agreement
  • rubric v4 · queue CIT-01

What we do

We grade your model's work against a rubric your team signs off on, and report a measured defect rate every week — on the queues where a confidently wrong answer is a legal, safety or revenue event.

Where we work

Judgment-heavy queues, not throughput.

Legal AI operations

Citation audit on demand packages, medical chronology accuracy, damages and valuation QA.

Marketplace and catalog

Attribute validation against pack imagery, allergen and dietary claims, substitution preference data.

Restaurant and delivery

Menu and modifier parity across channels, voice-order transcript grading, order-failure triage.

Model evaluation and preference data

Response grading, red-team and safety queues, ranked preference pairs for training.

How we start

A 30-day paid design sprint on one queue. Four artifacts you keep, whether or not the engagement continues.

Scoring rubric

Pass/fail defined per defect class, agreed with your team.

Gold set

150–300 verified examples. Versioned, and yours.

Defect taxonomy

Named failure modes, ranked by frequency and consequence.

Throughput and cost

A unit cost you can budget against.

The report you receive

Every period, on every queue.

Verified defect rate
0.42%
target 0.60% · prior period 0.71%
1,284 sampled · 94.1% agreement
actualtarget

Illustrative figures. Your numbers are whatever we measure.