Codex Astra did not remove the need for proof.
A more capable coding agent can cross a larger problem in one run. That makes scope, receipts, and independent verification more important, not less.
What happened, what we learned, and what changed. Stories from the space between an idea and something that works.
A more capable coding agent can cross a larger problem in one run. That makes scope, receipts, and independent verification more important, not less.
A beautiful knowledge graph becomes another inbox unless it is clear what belongs there, what stays elsewhere, and which system is allowed to call work complete.
The environment had terrain, trees, bunkers, and mountains. It still did not feel like the concept, because technical completeness and visual authorship are different jobs.
Before there was a playable hole, there had to be a way to keep concepts, assets, physics, builds, and claims from collapsing into one folder called progress.
The Mac app was telling the truth about paired capture while hiding the working single-camera path below the preview.
Moving a golf video off the iPhone was easy. Proving that the phone copy could be safely removed was the real system.
Swing Trainer's Live Studio passed its presentation suite, but a real 1149×772 window clipped critical capture controls. Why visual verification remains a separate release gate.
A signed macOS process, green tests, and an idle run loop did not prove that the user had a usable app. The launch investigation that separated disk work, window state, and observation error.
Swing Trainer needed to preserve partial human review without calling incomplete work training truth. The revision model that keeps drafts, labels, and source video separate.
Body motion can start and stop a golf recording, but it cannot prove where the club is. The evidence boundary behind Swing Trainer's hands-free capture system.
The physics boundary inside Swing Trainer: what monocular video can support, what calibration might unlock, and what still requires an external instrument.
Why the Swing Trainer RunPod gate stops at provenance, frozen splits, and real five-point labels—even with thousands of images already staged.
GitHub, Roboflow, Kaggle, and Hugging Face produced plenty of files. The hard part was refusing to call mismatched labels a dataset.
The club looked simple until motion blur, hands, clothing, and the edge of the frame turned one object into five different evidence problems.
The useful question is not “can Airtable automate this?” It is “what may this automation change without a person reviewing it?”
A pose system can see a body, a club detector can see a club, and a launch monitor can report a ball. Calling all three one measurement is how a golf product starts inventing facts.
The Range spike has a full instrument, a protocol, and a hard exit bar. It still has no iPhone measurement. That is not a gap in the story. It is the story.
We had no measured data and a simulator we needed to trust. A model-vs-model audit produced decision-grade findings anyway. The five moves, in order, and what this method can never tell you.
Most analysis sessions return confident prose. This one returned reproduction scripts, error bars, and a ranked list of experiments. The difference was a one-page brief written like a lab contract.
We had a tens-of-millions-of-throws simulator and an overnight Monte-Carlo run nearly three times larger. One Saturday-night physics audit showed the two models agree on every trajectory and disagree on half the outcomes — and every max-score throw was an artifact.
Why a markdown vault outperforms a Notion plus Asana plus Drive plus Slack stack when AI agents are part of the work. Battle-tested across QC, MHG, and Parley.
Twenty-six postmortems across QC and Parley in six weeks. The template, four worked examples, and the discipline that makes the same mistake stop happening twice.
A 48-contract system born from agents that praised their own work. The fix wasn't a smarter model. It was structural.
A Parley notebook reported three landmark architectures as broken on cross-signer ASL. A warmup and a gradient clip brought all three back, and two matched the best model. The ranking had measured my training recipe, not the models.
We trained seven landmark architectures three times each. Three of them worked on one seed and collapsed to near-random on the others. A single-seed comparison would have called two of those collapses a result.
Our sign model averages 42% across signers. That average hides a range from 26% to 64% — and the thing that decides where a person lands is not the signs they make, it is who they are.
Our best landmark-only sign model scores 45% on signers it has never seen. The field routinely reports numbers twice that high. The lower number is the honest one, and it is the one we publish.
I keep a running catalog of how hearing-led sign-language AI fails. It is not a list of other people's sins. It exists so Parley can catch itself the moment it starts to look like one of them.
Field notes from an operator running three ventures on Claude Code. Biweekly. No theory. Receipts only.
One AR-glasses decision for two ventures. The criteria that survived the cut, and why doing it once was the right move.
I started a Kaggle research project in Q2 2026 while running two startups. The decompression channel, the four open questions, and what makes it survive.
Claude is excellent at writing ESP32 firmware that compiles. It is not reliable at predicting what that firmware will do when the hardware is actually in front of you. Three incidents and the gate I added.
The answer isn't obvious. Systematic grid search gives you coverage. Random sampling gives you reality. You need both, and you need to know what each one is telling you.
A parametric physics engine, 8 parameters, two overnight runs. The database exists. Here's the honest accounting of what building it took and what we learned that we couldn't have learned any other way.
AR on physical objects is 80% coordinate system problems. Claude is great at helping you think through the geometry. You still have to understand it yourself.
Clean data beats model size. Every time. Don't upgrade the model until you've audited the labels.
LLMs in real-time hardware aren't ChatGPT. Latency budget is the constraint that changes everything.
Agents are unreliable judges of their own work. Here's how a structural fix — not a smarter model — stopped QC's CV pipeline from shipping silent failures.
Sixteen cookies that together are my whole Google account, dropped into a chat window — while building a security playbook. The how is the whole point.
Twelve-camera tracking rigs and Hawk-Eye are infrastructure at the top. The same capability now fits on a $249 board and a commodity camera. What that unlocks across every sport — and why the smartest way in is the narrowest one.
The routing model that holds up under load. Built for TruPath, tested on Mile High Golf, Quantum Caddy, and Parley.
A 20-agent Electron app that didn't ship. The decision to throw it away. What I'd tell anyone tempted to build the same thing.