AI evals as the highest-ROI skill for product builders
by @lennyrachitsky
ABOUT THIS BRAIN
Evals are systematic measurements that turn messy LLM interactions into actionable product improvements. They combine classic data-science error analysis with new LLM-specific tooling.
TECHNIQUES
KEY PRINCIPLES (13)
The goal is not to do evals perfectly, it's to actionably improve your product.
Perfectionism stalls progress; instead focus on rapid signal generation that guides iteration.
Why: Stochastic systems like LLMs have too much surface area for exhaustive coverage; actionable insight beats theoretical completeness.
"SPEAKER_3: The goal is not to do evals perfectly, it's to actionably improve your product."
Evals are the highest ROI activity you can engage in.
Teams that start error analysis get addicted to the insights and immediately see product improvements.
Why: A single systematic pass surfaces unknown failure modes that would otherwise stay hidden in production logs.
"SPEAKER_2: It's the highest ROI activity you can engage in. This process is a lot of fun. Everyone that does this immediately gets addicted to it."
Start with data analysis, not tests.
Jumping straight to writing evals skips the crucial step of understanding what actually needs measuring.
Why: LLM applications have unpredictable failure modes; you can’t design good evals without first seeing real user traces.
"SPEAKER_2: Usually, that's not what you want to do. You should start with some kind of data analysis to ground what you should even test."
Use a benevolent dictator for open coding.
One domain expert with trusted taste should label traces instead of a committee.
Why: Committees make the process expensive and slow; a single expert reduces cost while preserving accuracy.
"SPEAKER_2: You don't want to make this process so expensive that you can't do it. You can appoint one person whose taste that you trust."
Stop open coding when you reach theoretical saturation.
Continue reviewing traces only until new notes no longer reveal novel failure modes.
Why: Beyond saturation, marginal insights diminish; time is better spent fixing surfaced issues.
"SPEAKER_3: When you do all of these processes of looking at your data, when do you stop, it's when you are theoretically saturating, or you're not uncovering any new types of notes."
Never let an LLM do open coding for you.
LLMs lack the product context to spot subtle business-logic errors.
Why: Contextual knowledge (e.g., whether virtual tours actually exist) is required to judge correctness.
"SPEAKER_3: If I put that in a chat GPT and asked is there an error, it would say no, did a great job."
Use LLMs to synthesize axial codes from open codes.
After human labeling, LLMs can cluster similar notes into actionable failure categories.
Why: Synthesis is a deterministic summarization task that plays to LLM strengths.
"SPEAKER_2: You can categorize them with an LLM... Claude went ahead and analyzed the CSV file... came up with a bunch of axial codes."
Judge prompts must be binary and tightly scoped.
Ask for pass/fail on one specific failure mode rather than multi-point scales.
Why: Binary decisions are unambiguous and align with product requirements; scales drift in meaning over time.
"SPEAKER_2: You need to make a decision. Is this good enough or not? Yes or no?"
WHAT YOU GET
This brain captures how an expert actually thinks. Your AI retrieves their decision principles semantically and applies their reasoning to your situation.
Use this brain with your AI · OpenClaw · Claude · ChatGPT
principles · semantic retrieval · per-use pricing
Free during beta · Pay per use soon