Your first CLAUDE.md
The instruction file your AI reads on every session. What goes in it, what doesn't.
Lessons in AI-assisted building, from the first instruction file to the systems that keep complex work on track.
The instruction file your AI reads on every session. What goes in it, what doesn't.
Define a specialist. Scope, triggers, authority, tools. One file, one role.
What a skill is, when to write your own, the bootstrap pattern.
Five fields that turn fuzzy asks into concrete renders.
Voice, structure, source-pinning. The difference between a brief and a fight.
Inputs → research → structured output. The reporting loop in plain language.
Project scaffold, dev server, iterate-by-screenshot. End-to-end.
Specific language, used at specific moments, that produces outsized improvement.
Three models, three jobs. The decision rule that saves $300/month.
When to drop into Plan, how to read its output, when to exit.
The trust-but-verify reflex. Spot fabricated results before they ship.
The 6 slash commands worth memorizing on day one.
Long jobs offloaded right; the failure modes nobody warns about.
The 4-pass diff review that catches what tests miss.
The 20-minute setup, the gotchas, when MCP beats writing a tool.
Build CIPHER (legal/IP) from scratch in one sitting.
Allow vs ask vs deny — the model that prevents 90% of "wait I didn't mean that" moments.
The diagnostic that catches every kind of "finished but not really." One prompt, three patterns, one habit.
The substitution test for structural causes. Action items that ship. The close-out protocol.
MEMORY.md as an index. Four memory types. The single rule that prevents bloat.
SessionStart, PreToolUse, PostToolUse. Why hooks beat "agents must remember to."
How to add work to a contract without losing rigor. When a re-sign is required.
Communication protocols, delegation templates, authority boundaries.
Numbered folders. Single source of truth per concern. Why a markdown vault beats a SaaS stack.
Karpathy-style ingest. Page format, index pattern, lint cadence.
Frontmatter, options table, revisit conditions. Future-you reads this.
Why six is right and fourteen is too many. Authority levels, trigger words, escalation.
The contract format with three real examples — QC firmware, Parley research, MHG site search.
When worktrees beat single-branch. The merge etiquette nobody taught you.
The spend dashboard, the kill-switch, the weekly review that prevents bill shock.
The 4-question test before you spawn a subagent.
TDD with an LLM that wants to skip it. The 3 enforcement patterns that work.
Plan mode for senior operators. When to skip Plan, when to triple it, the exit ritual.
The review prompt that catches what humans miss. The escalation rules.
The 7-section postmortem template + how to make agents write it for you.
Why decision logs beat status meetings. Format, cadence, retrieval pattern.
Pitch without an NDA, publish a write-up, prompt a consumer LLM — each can be a disclosure. The grace period, the foreign-filing trap, and what to check before anything goes public.
Permissive vs strong-copyleft, the AGPL SaaS loophole, and why a single library can turn your whole codebase open. The license check that belongs in your workflow.
The hardest version of a CV problem is an open field; the easiest is fixed geometry you can pin with homography. How to choose the narrow, solvable entry point — then generalize.
A parallel technical project that restores focus to the main work instead of competing for it. How to tell if yours qualifies.
Picking a topic close enough to compound with the main work, far enough to feel like rest.
Declare in writing, before starting, that the arm will not become a product. The declaration is what makes it survive.
Three conditions under which a second project is a distraction, not an arm. Honest about when I would have said no.
Naming what you commit to NOT building, in writing, before starting. The Parley scope-rails block as the worked example.
Why monthly beats weekly and quarterly for a side project. The forcing function logic.
Using a public platform as the forcing function instead of a private repo. Why public is the discipline lever.
The advance commitments that make starting the arm psychologically safe. The three triggers that put it on ice.
Where to start when there's no CLAUDE.md, no contracts, no postmortems, and the deadline was yesterday.
When 6 agents are right, when 1 is enough, when more agents make you slower. With numbers.
Cache health ratios. Tokens per shipped artifact. Idle days vs. uncommitted-output days.
How the three disciplines compound. Where each catches the others' failures.
The amendment protocol at scale. Refusing 'while we're in there' work. Without becoming an obstacle.
Cross-venture routing, shared agents, isolation rules. Running 3+ ventures from one operator.
Pod lifecycle decoupling. Exit-path audits. Watchdog cadences. Don't lose a 14-hour run.
The signs that your session is gradually losing accuracy. Restart hygiene that compounds.
When to build vs use existing. The protocol in plain language. The failure modes that cost teams two weeks each.
Beyond chief-of-staff routing: parallel-spawn, gather-then-merge, race-then-cancel. When each fires and when each breaks.
When an agent misbehaves in prod: the diagnostic ladder. Five rungs from "claim says done but isn't" through context drift, tool failures, schema mismatches, upstream model regressions.
The 5-section agent file pattern. Two case studies from the TruPath team.
How to ship Phase-0 R&D against unreleased hardware (Maverick AI).
The on-call playbook: paging, evidence collection, triage ladder, and postmortem-on-rails.
The audit-trail format regulators (and your future self) will thank you for.
Hooks + statusline + a kill-switch script. The whole stack.
The 12-document SBA 7(a) packet, agent-assisted, with rework rate measured.
Building an ASR eval set from scratch when there's no public benchmark.
Two models that disagree tell you almost nothing until you can regenerate the one you are auditing. Reproduction converts ambiguous divergence into a named mechanism, and it finds defects before the correction work even starts.
Two models agreed on every trajectory to a hundredth of a foot and still flipped 48.3 percent of scored outcomes. Simulated states degrade gracefully; simulated labels fail all at once, right where the outcomes get interesting.
Friction swung our prediction 14.9 inches of a roughly 15-inch budget, three times the runner-up, and mass was noise. A one-afternoon sensitivity sweep turned the measurement wishlist into a spending plan.
The audit put ±2 inches of uncertainty on our flight prediction and ±15 on everything after contact. The failure lived in one stage, and a per-stage budget is what kept us from throwing away the stage that worked.
The QC physics audit produced runnable models, a dozen figures, and a stack of CSVs. The artifact that mattered three weeks later was none of them — it was a short brief written for the engineers who have to act on the findings.
We assumed spin mattered the way it matters on a golf ball. Priced in the same units, the borrowed mechanism was worth 3 inches and the real one 23, a 7× miss in which mechanism matters, caught in one afternoon before it spent our measurement budget.