Lab Notebook · Experiment Nº 001
Shipping an AI feature without evals is shipping on vibes. Build a real eval for your feature: a quality rubric, a starter golden dataset, and a pipeline your team can run with. Free, no signup, runs in your browser.
A quality rubric
The dimensions that matter, weighted, with hard-fail criteria
A golden dataset
30 to 50 test cases across happy path, edge, and adversarial
An eval pipeline
Automated checks, LLM-as-judge, and human review cadence
Pass thresholds
A bar derived from stakes and reversibility, with the math shown
You'll walk out with a quality rubric, a starter golden dataset, an eval pipeline recommendation, and a complete plan you can hand to your team.
1. Describe your feature
What it does, what good and bad look like, what a failure costs you
2. Quality rubric
AI drafts scoring dimensions from your description; you edit anchors, weights, and hard-fails
3. Golden dataset
AI generates starter cases across happy path, edge, and adversarial; you curate
4. Eval method
Automated checks, LLM-as-judge setup, human review cadence, tradeoff boundaries
Your work saves in your browser as you go. What you type in Step 1 is sent to an AI model only to draft your rubric, dataset, and method, and is not stored server-side.
Optional background. The tool above stands on its own.
A plan is the easy part. I help teams stand up eval pipelines that actually run, and keep running after I leave.
Work with me