Kuan Yu
← Journal

15 Days with Fable

Fable 5 came out two weeks ago. This is one long article on everything I actually built with it: 15 builds, real numbers, live links where they exist, and the embarrassing parts left in. It grows as each build lands.

TECHNOLOGY AI BUSINESS ENTREPRENEURSHIP

Fable 5 came out two weeks ago. I’m a solo founder in Singapore, and I decided the honest way to show what this model changes is not a hot take. Receipts. So: 15 projects, one article, updated as each build lands, every build answering the same four questions. What it is. What I actually did. What you can do with it. Who it’s for.

”How’s Fable so far?”, honestly

Since everyone asks, here’s my actual answer before the receipts.

It may not be an Opus upgrade at all. A community similarity analysis (this one) found Fable doesn’t cluster tightly with Anthropic’s other models the way siblings usually do. The author’s own speculation: a new base model on a different data mix, a different animal rather than a bigger one. Speculation, flagged as such. But it matches the driver’s seat: it doesn’t fail where I’d calibrated for failure.

The real value for me is unknown unknowns. People keep saying the models are really good. But if you can’t accomplish your goals with them, how good are they really? My pattern for two years: stuck, and either not knowing why, or knowing why and still not getting past it. Every model before gave me a partial diagnosis: “your problem is selling,” “your problem is how you build.” Partial diagnoses feel true and change nothing. Fable’s read of everything I’d done was the first one complete enough that things actually started moving. Most of this series exists because of that.

“Just send it” was never the problem. Earlier models would draft something and tell me it’s ready. Just send it. But a send with my name on it has standards: is it up to my bar, is every claim verified, is it my voice, do I understand it deeply enough to defend it in the reply? A draft I can’t defend is not a send. What finally worked is a pattern I think of as the barrier breakdown: an artifact holding the message, every file it’s built on, why each one is the way it is, and the exact steps for me to verify and approve. A memo deep enough to give me real understanding in minutes. Then the send happens. That’s how a client send finally fired, how an application I’d sat on for weeks left the building, and how I shipped my first video demo ever. I don’t know how to make videos. It made the video; I reviewed it; it went out. (The public, copyable version of this idea is the firepage build below.)

The meta-move underneath it all I learned from Jordan Peterson years ago: when a problem is too big to solve, break it down until the pieces are too small to procrastinate on. These artifacts just mechanize that, at the exact moment of maximum avoidance.

Ops notes for builders: I run two Claude Code Max accounts in parallel now, plus Azure credits running GPT-5.6 for the bulk work. I use Fable to write the skills, harnesses, and instructions so cheaper models can operate closer to its level when it isn’t around, and to direct lower models (and Codex) rather than doing everything itself. And I tried the recurring-loop pattern everyone recommends; an audit of my own usage showed my work is mostly one-off diagnosis and fixes, not recurring pipelines, so I stopped forcing it. Know which kind of user you are.

Ground rules for the series: real numbers over adjectives, the failure log stays in, clients and communities stay anonymous, and every build ends with either something you can use or something you can ask me. If a build reads like a report card, I’ve failed. Tell me.

The builds

Ranked roughly by how much I think you’ll enjoy them. Each becomes a full section of this article as it lands; the list shows what’s live so far.

  1. Outbound Systems That Run Every Single Day: the flagship: 1,400+ individually personalized DMs, a verified 5% cold reply rate, and the official API that undercounted replies 7x. (live, below)
  2. My Feed Filter Admits When It Gets Worse: rai, the local-AI feed filter, and the eval score that dropped 16 points when the test got honest. (live, below)
  3. Seven Gates and a 95-Cent Receipt: the outreach engine you can run on any company, with a fully costed $0.947 run published. (live, below)
  4. He Already Fired a Human for This Job: a daily SGX ledger on a brutal deal: 20-30 days, zero misses, before it’s allowed to have a price. (coming)
  5. The Seller’s Bot Promised a Refund. The Human Disowned the Bot.: AI as family consumer-rights lawyer; full ¥2,199 recovered. (coming)
  6. “Live Mode Has Never Worked”: a 34-Hour Rewrite: a live call copilot: 0.04s captions, 37.5MB flat RAM, and the checkbox that’s still honestly unticked. (live, below)
  7. My Todo Board Argues Back: the AI operating system that runs my company of one and quotes my own constitution back at me. (live, below)
  8. Cite or Shut Up: One Pattern, Three Products: the centerpiece: 6,080 filings, 32 citations, 0 fabrications, in 34 minutes. Then the same spine on health data and live chat. (coming)
  9. 11 Months, 12 Repos, 0 Customers. Then One Day.: the tax-tool redemption arc; the answer key came first this time. (coming)
  10. The Plan Refuses to Render: the paid training-plan engine: 47 checks, a blind AI-vs-AI review, and the day it correctly refused me too. (coming)
  11. One Folder Reads 14 Inboxes: comms triage with a fresh-eyes critic and a hard invariant: no approval record, no send, ever. (live, below)
  12. The Draft Was Never the Bottleneck: firepage, one HTML file that gets stuck messages sent, extracted from a real multi-week failure. (live, below)
  13. Two Kinds of Waiting: the family medical-ops system, and the red/green split that stopped years from leaking. (coming)
  14. Small Tools, Hard Rules: a dividend screener that caught a fake 35% yield, and a watcher whose busiest code path is its own safety boundary. (live, below)
  15. My Running Coach Argues With a 1998 Textbook: VDOT verified against Daniels to within seconds, dew point priced into your pace, and the race it just coached. (live, below)

The scoreboard

I’m tracking this series in public: reads (site analytics), repo traffic (GitHub clones/views, baselined the day before this article went up, effectively zero), and the number that actually counts on my board: paying customers using something weekly. The series doesn’t move that number until a stranger uses or buys something. That’s the honest frame: this article went up on day 27 of a 100-day public build, and the scoreboard updates weekly.


Build 1: Outbound Systems That Run Every Single Day

Cold outreach as an operations discipline: the machine sends, the humans converse, and nothing gets to self-report.

Cold outreach that works is not a growth hack. It is an operations discipline: a system that runs daily without supervision, measures itself truthfully, and puts a human in charge of every conversation. I design, build, and operate these systems end to end. Here is how I ran a recent one, and the standards I hold it to.

The operation

Over 65 days I ran a fully automated, individually personalized cold outreach campaign on X: 1,400+ DMs, a verified ~5% reply rate, a permanent log of every message sent, and not a single message ever sent to a human by a machine acting alone. The pipeline behind it: a 3,000-row scored prospect list built from public profiles, messages drafted one person at a time, daily volume that paces itself, and a human-operated reply desk.

For a fully cold channel, a verified 5% reply rate is the headline. Getting it, and being able to prove it, took three disciplines most outreach setups skip.

Discipline 1: it runs every day, and I know it ran

Outreach compounds only when it never stops. Small, steady daily volume with respectful pacing beats bursts every time, so the system computes its own daily budget from its own history: volume grows on clean streaks and adjusts down the moment the data says to. Pacing is randomized and spaced the way a careful human would work. The same rule I apply to everything: volume may match measured history, never exceed it.

Reliability is designed in, not hoped for. The system reports in daily with a summary that reads in 30 seconds, and the absence of that summary is itself the loudest alarm: a dead system cannot announce its own death, so the alert fires on silence. If anything interrupts a day, the day continues where it stopped instead of starting over, and the prospect queue can never quietly run empty.

Discipline 2: measure what actually happened, not what an API claims

The platform’s official API turned out to be blind to most real replies: on one account’s back catalog it reported 6 replies where 41 actually existed, a 7x undercount, because it does not understand the platform’s newer conversation format. Most operators would have read the API’s number, concluded the campaign was failing, and killed messaging that was working.

I caught it because nothing in my systems gets to self-report. Every reply is verified against what a human eye would actually see on the page, automatically, thread by thread, five minutes a day. Every message sent is a permanent logged record with a timestamp and an outcome. When a dashboard disagrees with reality, the dashboard is treated as the bug: in one review session an operator flagged three numbers that looked wrong, and all three were found and fixed the same day.

The same rigor applies to the AI layer. Before any change to how the AI scores prospects or drafts messages goes live, it has to beat the current version in a measured test. If it only ties, it does not ship. That is why the numbers mean the same thing week after week.

Discipline 3: machines draft, humans converse

Message quality ran as controlled experiments, not opinions: no variant judged before 150 sends, anything under 2% killed, champion and challenger always running. The data taught us things templates never could: the best opener style outperformed a generic one by 1.7x, short messages under 400 characters received exactly zero replies, and targeting quality moved results roughly four times more than wording. Every message was drafted individually against the specific person’s public profile, then checked by gates before it could send.

And the hard rule that makes the whole thing trustworthy: no machine ever replies to a human on its own. Replies are drafted from a playbook built out of the questions real prospects actually asked, pass three content gates, and a human approves every single send. Automation earns its keep by making the human fast, not by replacing their judgment.

The standards, in short

  1. Every message sent is a permanent, timestamped record.
  2. Every reply is verified against what a human would actually see.
  3. Daily volume is earned from track record, paced like a careful human.
  4. Silence is the alarm; reliability is designed, not hoped for.
  5. AI changes go live only after beating the current version in a measured test.
  6. Humans hold every conversation.

Who it’s for

Anyone running automated outreach and guessing at invisible limits; builders who trust platform APIs to report their own numbers; and anyone whose growth project has beautiful telemetry on everything except whether it is working.

If you have a product people already love and you want outbound run against it with this level of rigor, personalized at the individual level, measured truthfully, human on every reply: message me at @phuakuanyu. I will walk you through live logs and dashboards from a real operation before you commit to anything.


Build 2: My Feed Filter Admits When It Gets Worse

A local AI that curates my feeds, and the eval score that dropped 16 points the day the test got honest.

Why it exists

I scroll X for work. Signal is real there: papers, launches, postmortems, people thinking in public. So is the sludge engineered to sit between me and it: engagement bait, crypto pumps, thirst traps, the fortieth variation of the same hot take. Attention is the only budget that matters in my company of one, and I was spending it like a tourist.

rai is my third attempt at fixing this. The first shipped in December and found two users. The second died mid-rebrand in a half-finished cloud rewrite. The third started over in March with a different premise: run the model on my own machine, and measure everything.

What it is

A Chrome extension that filters X and YouTube with a small AI model running locally through Ollama. Nothing about my feed leaves my computer by default. The pipeline is three sieves:

  • A regex prefilter kills the obvious junk instantly, eleven categories of it, no model needed.
  • Whatever survives goes to the local model, which scores each post on four dimensions: novelty, specificity, density, authenticity. Below threshold, hidden.
  • Images and video thumbnails get one more gate: a local vision model screens for bait imagery before any blur lifts.

Everything is CSS on the already-rendered page. No platform APIs, no auto-anything; even reply suggestions are copy-paste only. If Ollama is not running, the filter fails open and the feed is just a feed. And for people who will never run a terminal, a cloud mode swaps the local call for a hosted one with a one-click free key: no account, no email, no card.

YouTube gets the treatment I actually needed: I open it to play music and resurface forty minutes later, three Shorts deep. So rai blurs the entire home grid by default and reveals only music. After ten Shorts or five minutes, it interrupts with the only question that matters: is this really where the evening goes.

The eval that got honest

Here is the part this section is named after.

Every prompt or model change in rai has to pass a golden-set eval before it ships. That is the whole differentiator; feed filters normally run on vibes. In early July my golden set was 109 items, and the production model scored beautifully against it: 100 percent critical-signal recall, 96.2 percent signal, 89.3 percent noise caught. Those numbers went on the website. They were true.

Then I rebuilt the test from reality. Three days of my own live traffic, 5,309 logged decisions, 1,058 of them machine-judged, every critical call hand audited, distilled into a harder 294-item set with tiers the old test never had, including the category that actually matters: tempting noise, the stuff engineered to look like signal. Same production model, same prompts, new test: 90.9 critical, 80.1 signal, 60.8 noise. Signal recall fell 16 points overnight. The filter did not get worse. The measurement got honest, and honest measurement is the entire product.

So the drop is published, here and on the site, replacing the prettier numbers. A filter that cannot admit it got worse cannot be trusted to say it got better. The same is true of the people who build them.

The big model hires the small ones

The coolest part, to me, is who runs the hiring. I am a Claude person; this whole series is me building with a frontier model. But rai itself runs no frontier model at all. Everything that touches my feed is a small open model on my own machine.

The frontier model’s job is one level up. When new local models land on Ollama, Fable reads what is available against what this project actually needs, writes the tests, pulls the candidates, and runs the bake-off on my machine: accuracy against the golden set, speed, everything, then keeps the winner and deletes the rest. The last sweep benchmarked twelve local models in one pass, phi4-mini, gemma2, a stack of qwen3 variants, granite4, Nemotron, MiniCPM and more, against the production default. The incumbent kept its seat, which is exactly what a real gate looks like: most challengers should lose.

The same loop maintains the tests themselves. It watches what I actually browse, logs the decisions, and periodically rewrites the evals, the verification, and the prompts from that record. The 294-item golden set that cost me 16 points of vanity was written this way. So the division of labor comes out clean: the small models do the daily work in private, and the big model shows up as the staffing agency, the examiner, and the prompt writer. Each model where it is strongest, and my feed never leaves my machine.

The other honest numbers

My Chrome Web Store listing says 2 users. Both are inherited from the December version’s listing; net new adoption of rai itself is currently zero verified humans. I also built a tips ledger that captures workflow ideas spotted while scrolling: 110 captured so far, 0 triaged. The capture loop works and the follow-through loop does not exist yet, which is a sentence that describes more of my tooling than I would like.

What you can steal

  1. Gate every prompt change on a versioned eval, and keep the eval where you cannot quietly ignore it. A regression gate you can bypass is a mood.
  2. Rebuild your test set from real traffic, not synthetic examples. The score will drop. That drop is information you were previously hiding from yourself.
  3. Fail open. A filter that breaks should give you your feed back, not take your feed away.
  4. Local by default, cloud by choice. Privacy is an architecture, not a settings toggle bolted on afterward.

Who it’s for

People who scroll X for work and lose the thread to engineered noise; people who open YouTube for one song; anyone who would rather run a small model on their own machine than mail their attention patterns to a cloud. Install from the Chrome Web Store in cloud mode, or go full local: github.com/phuaky/xrai.


Build 3: Seven Gates and a 95-Cent Receipt

A company name goes in. A personalized cold email comes out. Seven gates in between, and a human holding the only send button.

Why it exists

Drafting a cold email with AI is trivial. That was never the problem. The problem is the last step: sending is an external, irreversible action, done in my name, to a stranger I might want as a customer. An email that is generic, wrong, or embarrassing does not just fail; it burns the lead and my domain reputation with it. So the real design question was never “can AI write outreach?” It was: how many checks does a machine need to pass before I would let it near a real inbox?

This build is my answer, and it is the coolest thing I have designed. Not because any single part is clever, but because the whole thing is one connected machine: research, judgment, drafting, gates, cost accounting, and a human approval that nothing can route around.

How it works: The Line

Every lead rides a 15-station pipeline I render publicly as “The Line.” The spine of it:

Scout researches the company from public sources. ICP scores the fit 0-100 across explicit dimensions and kills bad fits on the spot; in one real campaign, 35 of 90 leads died right here, a 39% kill rate at the first real checkpoint. Strategist picks the angle, a judgment call an LLM now makes instead of the keyword matching I started with. Personas maps who decides and who blocks. Artisan builds the free value artifact, the thing the email gives away rather than asks for. Poet drafts the email around it.

Then the gates, seven of them, in order:

  1. ICP Match: is this still the right company for what we sell?
  2. Critic: an adversarial AI pass over the draft’s claims and logic.
  3. Deliverability: will infrastructure even let this arrive?
  4. Spam Score: does it smell like spam to the filters?
  5. Format: plain text only, structure rules enforced deterministically.
  6. Voice: does it sound like me, not like an AI mail merge?
  7. AI Judge: the final quality bar, enforcing a value-first rubric: the email must give before it asks.

Four of those are deterministic code that costs $0 and runs in 0.02 seconds. Three are AI passes that cost real money, and that asymmetry is the design: cheap hard rules kill the obvious failures so the expensive judgment only spends on drafts worth judging.

And after all seven: a human. The send button is not in the pipeline. Drafts render in a read-only review UI, and approval is a deliberate separate act. That is not caution theater. Before the approval gate existed, the system once sent an email it should not have; it caught and quarantined the mistake itself within the hour, and the structural gate exists because of that day. The rule since: machines can do everything except the one thing that cannot be undone.

The receipt

I published one full run, timestamps and costs real, business name changed: $0.947 total to take one lead through all 15 stations. Artisan $0.296 over 210 seconds. Critic $0.124. Poet $0.265. AI Judge $0.263. The four deterministic gates, $0.00. Token flow: 141,500 in, 39,800 out.

Under a dollar for research, a scored fit decision, a custom value artifact, an adversarially reviewed draft, and a judged final email. That is the number that makes the whole category real: the marginal cost of doing outreach properly is now lower than the cost of doing it lazily ever was.

The engineering sprint

The system is 36 sub-workflows, roughly 295 discrete steps, governed by a spec with 108 verifiable criteria, covered by 534 passing tests. The two weeks that finished it are the most fun I have had building anything:

  • Five features in parallel, merged the same day. On July 11 I ran five Claude Code agents in five separate git worktrees: the ICP gate rework, the LLM strategist, the Personas station, a bring-your-own-product field, and a second model provider wired in behind the same agent abstraction. All five merged to main that evening. 41 commits that day, solo.
  • Spec to live public deploy in under 24 hours. Cloudflare Worker + KV intake, a Pages site, a local worker daemon, and then three production-only bugs found and fixed by actually driving the deployed thing in a browser: a dead model id silently killing every Scout call, a JSON rendering bug, and a CORS-plus-redirect pair that only existed on the real domain.
  • Prompt-injection hardening on day one, not as a later pass. The public demo accepts free-text company names from strangers, so the agent behind it runs on an explicit tool allowlist. A hostile input can waste one run’s tokens; it cannot touch a shell or a file.

The eval that caught its own coin flip

I built a frozen 16-case eval harness that live-replays real historical inputs through the real Poet and Judge. Same prompts, same model, same config, runs a day apart: 9 of 16 one day, 0 of 16 the next. Nothing changed except the judge’s mood. A same-day model swap scored opus at 1 of 16 against sonnet’s 9 under the identical scaffold.

That variance is unresolved, and I am publishing it anyway, because it is the most useful thing the harness has produced: an LLM judge is not a test suite. “It passed the benchmark” and “the tests are green” are different claims, and any builder shipping LLM-judged pipelines is sitting on this same coin flip whether they have measured it or not. Multi-run averaging is the obvious next step; single-run scores are theater.

Run it yourself

The demo is live: type any company’s name, optionally what you sell, and watch The Line process it station by station. No account, no key. Ten real runs sit in the gallery, all named public companies, ICP scores 45 to 88. Nine passed the Judge. One, a beloved bakery chain, scored 45 and renders as an honest “here’s why I’d skip it” instead of being hidden. Three more runs were Judge-rejected and deliberately kept out of the gallery, because a quality gate you never see fail is indistinguishable from marketing.

The honest limits, also published on the page: live generation runs on my own Mac, so a run can take a while if it is asleep; the demo is rate-limited; and it can never find or contact a real person. The Enricher is disabled at the type level in demo mode. Drafts address a role, never a scraped name. The console even lists its own five unresolved drift issues. If I am selling verification, the demo has to survive being verified.

What you can steal

  1. Gate asymmetry. Deterministic $0 gates first, expensive AI judgment last. Never spend model money on a draft a regex could have killed.
  2. Make the irreversible step structural, not procedural. No flag, no config: the send path physically requires a human act. Design for the incident you already had, not the one you imagine.
  3. Publish the receipt. Per-stage cost and time for one real run buys more trust than any capability claim.
  4. Show a failing gate. Keep a rejected run visible. A gate that never visibly fails is not credible.
  5. Live-replay evals, then distrust single runs. Replaying real inputs through real code found what mocks never would: the judge itself is noisy.

Who it’s for

Builders deciding whether a multi-agent pipeline can be trusted with an external, irreversible action; GTM people who want to see the actual research-to-draft trail instead of a claim; and anyone curious what a solo 96-commit sprint with parallel agents actually produces.

Run it on your company: gtm-demo-9h0.pages.dev/demo/. The annotated, fully costed run is at /story/.


Build 6: “Live Mode Has Never Worked”: a 34-Hour Rewrite

A live call copilot for people who lose the thread mid-call. Plus the checkbox that is still, honestly, unticked.

The problem: a live call is three jobs at once

If you sell what you build, the discovery call is where everything is decided, and it is quietly one of the hardest things a founder does. You are doing three jobs simultaneously: listening to what the person is actually saying, steering the conversation through some structure, and deciding in real time what to ask next. Do two of those well and the third collapses. I would come off calls having talked too much, having skipped the one question that mattered, holding half a memory of what they actually said.

The existing fixes all fail in the same two ways. Meeting bots join your call as a visible third participant, which changes the conversation the moment it appears. And almost all of them do their real work after the call: a summary, a transcript, a coaching note, delivered exactly when it can no longer help you. The call is live. The help arrives dead.

What it does for me

call-copilot is a Mac app that listens to the call I am already on, no bot in the room, and does the third job so I can do the other two:

  • Both sides transcribed as we talk. Their words and mine, captioned live, so nothing said in minute four is lost by minute thirty.
  • A dashboard that knows where we are. It tracks the phase of the conversation, keeps a running fit verdict, and maintains an intel log of every fact learned so far, growing as the call goes.
  • One line telling me what to ask next. Not a script. A prompt, in context, at the moment I would otherwise flounder.
  • My methodology, not a generic one. The coaching logic lives in three swappable markdown files, Mom Test, Hormozi’s CLOSER, a generic discovery flow. Changing how I sell is editing prose, not code.

And the property that makes it usable for real conversations at all: nothing leaves my machine. The transcription model runs locally, the audio stays on the laptop, and no third party attends my customer’s candid moments.

The result is a different kind of presence. I stop performing memory and structure, because the machine holds both, and I get to actually listen. That is the product: not notes about the call afterward, but attention during it.

The 34 hours

The embarrassing part first: I built v1 in the spring, and for three months my own engineering doc contained the sentence “Live mode has never worked.” The only working mode replayed yesterday’s calls, slower than real time. Five small plumbing wrongs added up to a tool that missed its entire point.

In mid-July, one 34-hour window rewrote it: continuous two-stream capture, a transcription server that loads the model once, and a thinking loop decoupled from the listening loop so the app is never deaf while it reasons. The measurements, from the verification log, same night:

  • Utterance to caption: 0.04 seconds average, 0.14 worst case.
  • A twenty-minute soak on a real call replayed as live audio: 40 thinking cycles, zero crashes, memory flat at 37.5MB.
  • Sixty seconds of silence: zero transcript lines. The old pipeline used to hallucinate the word “you” out of dead air.
  • Kill the brain process mid-call and the captions keep flowing while the dashboard holds its last good state. 18 of 24 verification criteria ticked.

The unticked checkbox

Which brings up the six that are not. Everything above was proven on a real call replayed as if live: honest audio, honest work, simulated liveness. The four criteria that certify a true live call, captured through the operating system while a real human talks over you, are still marked pending in the log. I could have quietly run that test before publishing this and ticked the box. I have not yet, and the log says so. A verification trail you backfill for the blog post is not a verification trail. The claim stands at exactly its evidence: solid on replayed reality, unproven on live reality.

What you can steal

  1. Aim tools at the moment of need, not after it. A transcript tomorrow is worth a fraction of one line of help mid-call.
  2. Keep methodology in markdown. The part of a tool you iterate weekly should be editable at the speed of prose.
  3. Local-first is a feature for exactly the conversations that matter most. The calls you most want help with are the ones you least want uploaded.
  4. Publish the unticked boxes. They are the reason the ticked ones mean something.

Who it’s for

Founders and solo operators who run their own discovery and sales calls on a Mac and want live help without a bot in the room or audio in a cloud. It is a builder’s tool today; productizing it waits for a real buyer, per my own constitution. If that buyer is you, the conversation costs nothing: @phuakuanyu.


Build 7: My Todo Board Argues Back

The system that planned, argued about, and then built this very series.

Why it exists

My Claude Code setup right now: two Claude Max accounts in parallel, plus Azure credits running GPT-5.6 for the bulk work. For practical purposes, the token ceiling is gone.

Everyone assumes that’s the dream. It’s where the real problem starts. When execution is unlimited, the binding question flips from “can I build this?” (always yes now) to “is this the thing worth building today?”, and the model cannot answer that one. It optimizes for the request in front of it, not the mission. Left alone with infinite hands, I will research forever, polish what nobody asked for, and ship nothing to a stranger, all with perfect craftsmanship.

founder-home exists to answer the question the tokens can’t: what should today actually do?

What it is

Think: a chief of staff who has read your constitution, and holds you to it.

founder-home is the operating system for my company of one. Not a productivity app. A repo with three load-bearing parts:

  • One board. A single state.json holds today’s goals, the task stack, every open send, and the one metric that counts. Everything else reads from it. There is no second source of truth to hide in.
  • A ledger I don’t write. Every Claude Code session in every repo (nineteen and counting) writes a dated entry via hooks: what was asked, what was committed. The morning board reads yesterday back to me. I can’t misremember what I did, because I’m not the one keeping the record.
  • A constitution the AI enforces. My failure modes are written down and named. “The Loop”: research → build → abandon, twelve archived projects of evidence. “Vanity Theatre”: polishing the mirror instead of asking a stranger to pay. Rung priority: paying customers before prospects before interesting work. A NOT-DO list. And one binding constraint above everything: a send leaving the building outranks build, research, or harness work.

Every morning the AI reads the stack, proposes one needle goal and one support goal, debates me, and compiles the winners into structured prompts with a done-when and an explicit Anti line: the failure mode this goal must not become. Then it fires.

The part that actually matters

The board prints my worst number first, every day: COIN, paying customers using something weekly. On purpose. The system is designed so I can’t grade myself on effort, cleverness, or how good the tooling looks. This article you’re reading counts for nothing on that board until a stranger pays and comes back.

And it argues. The morning this series launched, I arrived with fresh energy and a new plan, one day after a discouraging customer signal. The system’s job in that moment was to say: this pattern has a name in your constitution, the sends still outrank it, here’s the goal restructured so the sends fire anyway. It did. They did. The new plan became this series. After the sends, not instead of them.

What you can steal

The pattern is portable and none of it is secret sauce:

  1. One state file. Any second board is a place to hide.
  2. Hook-written ledgers: your AI logs every session automatically; you never self-report.
  3. A constitution file with your failure modes named. Naming is what makes them catchable in the moment.
  4. An Anti line on every goal: what this must not turn into.
  5. One honest metric printed first, especially while it’s ugly.

The uncomfortable prerequisite: you have to write down your own failure modes, in your own words, and give the AI standing permission to quote them back at the worst possible moment. That permission is the product.

Who it’s for

Solo founders and builders whose AI produces endless work but no accountability; anyone whose productivity system has never once told them no.

I’ll publish a sanitized starter template (hooks + empty board + constitution skeleton) if enough people want it. Tell me: @phuakuanyu.


Build 11: One Folder Reads 14 Inboxes

Email triage as a conversation, a fresh-eyes critic on every draft, and one hard invariant: no approval record, no send, ever.

Why it exists

A confession first: I do not open email. Not “I am behind on email.” I do not open it. Fourteen email accounts across my businesses and projects, four live WhatsApp numbers beside them, and at one honest count, about 3,600 messages sitting across those inboxes with over 1,200 unread. One account alone held more than half. For a solo operator, every inbox is a room where someone might be waiting, and I had eighteen rooms I refused to enter.

The fix was not discipline. I have met myself. The fix was making the rooms come to me.

What it is

There is no product here, no dashboard startup, no inbox app. My mail already lives on my machine, and Claude Code can read it there. So the mail client is a conversation, and the loop is scheduled sessions of the same agent I build with all day:

Every morning at 07:30 a sweep reads all eighteen channels into one folder. It tiers what it finds: deterministic rules first, model judgment second. The noise, receipts, newsletters, notifications, gets categorized and marked as read, and I never see it. The real things get understood: who this person is, pulled from my own local contacts and history, what they are asking, what our last exchange said. At 07:45 the reply pipeline runs: for each message that deserves an answer, it assembles the context, states a goal and a stance, and writes the draft. Then the part I trust most: a second pass, from a fresh context that never saw the drafting, reads the draft cold and critiques it. At 08:00 my phone gets a one-glance summary. Two smaller scans at 13:00 and 18:00 keep the folder current.

What used to be eighteen rooms is now one two-minute read and a queue of drafts waiting for my yes.

The invariant

The rule that makes the whole thing safe to run unattended: no approval record, no send, ever. Every send requires my explicit, logged, per-item yes. And the honest version of today’s state goes further: even after my yes, the send itself is still a manual step. The automation reads, sorts, understands, and drafts. It has never once sent a message by itself, and the approval ledger is the proof. An email with my name on it is a promise, and promises do not get delegated to a scheduler.

The system also assumes it will fail, and fails loudly. 196 scheduled runs over the last 23 days: 177 succeeded, just over 90 percent. The other 19 texted me that they had failed, instead of going quiet. In a system whose entire job is “you never have to check,” silent failure is the one unforgivable bug. Some things did go quiet anyway: one of the fourteen accounts spent weeks failing its scan, and the WhatsApp bridges have a known blind spot the health check names instead of hiding. Both are on the board. A triage system that hid its own gaps would be the thing it replaced.

It learns, slowly and on the record: every time I correct a tier, the correction is logged and the rules absorb it. Three corrections so far. The folder is better at knowing what matters to me than it was three weeks ago.

What you can steal

  1. The pipeline shape: scan, deterministic rules before model judgment, context pack, stated goal and stance, draft, then a critic in a fresh context that never saw the drafting. That last split is the quality jump: writers cannot grade their own work, and neither can models.
  2. The invariant, word for word: no approval record, no send, ever.
  3. Failures must page you. The absence of a heartbeat is an alarm; a quiet crash in a trust system is worse than no system.
  4. Everything is a plain file: queues as JSON, logs as JSONL, config as markdown. Diffable, auditable, no database between you and your own state.

Who it’s for

Anyone running more than two or three real inboxes through one skull, especially solo operators with several fronts. And specifically: people like me, who were never going to open the rooms. The triage does not make you a better correspondent. It makes not-opening-email survivable, which on a bad week is the same thing.


Build 12: The Draft Was Never the Bottleneck

One HTML file that gets stuck messages sent. Built in 25 minutes, extracted from a failure that took weeks.

The problem nobody builds for

AI writes a perfect draft in three seconds. Everyone is building better drafting. But watch what actually happens to a finished draft: it sits. Mine sat. A high-stakes application, fully written, fully verified, ready in every way a document can be ready, stayed unsent through a multi-week drought while I re-read it and waited for a better moment that was never coming.

Here is the uncomfortable finding from my own logs: throwing more verification at a stuck send does not unstick it. I tried. A draft can be checked seven ways and still not leave the building, because the bottleneck was never confidence in the text. It is the send itself: an open-ended, witnessed, irreversible act that perfectionism can defer forever.

What finally fired that application was not a better draft. It was a page.

What it is

firepage is one self-contained HTML file. No server, no build step, no dependencies. Works offline. You fill four blocks:

  • a verification stamp: each claim in the message, checked, with its source
  • the draft itself, collapsed behind a copy button so you cannot re-read your way back into doubt
  • a table of what the message is built on and why each piece is the way it is
  • a finish checklist, six items or fewer

Then the mechanism that does the real work: the last checkbox is not “done.” The last checkbox is telling a specific, named person “fired.” Completion gets a witness and a receipt, not a feeling. A send you must announce to someone is a send that happens.

The numbers, small on purpose

Four commits, one sitting, 25 minutes from first commit to last. Three HTML files, between 114 and 165 lines each. A README and an MIT license. Zero stars, zero forks, because this article is the promotion and I baselined the counters the day before it went up. One real send unblocked by the pattern before it was ever generalized, which is the only number that matters.

The pattern earned its way here: it started as a private rule I now apply to every outbound message I write. When the same page shape had fired real sends enough times, I spent the 25 minutes making it copyable.

What you can steal

The whole thing, literally: github.com/phuaky/firepage, MIT, and the live example at phuaky.github.io/firepage. Copy template.html, fill the four blocks, open it in a browser.

But the portable idea is smaller than the file. When a task is too big to do, break it down until the pieces are too small to procrastinate on, and give the last piece a witness. The page just mechanizes that at the exact moment of maximum avoidance.

Who it’s for

Anyone who drafts fine and does not send. It assumes the draft already exists and it is not for cold outreach at scale. It is for the message you have been carrying for two weeks, the one with a real person at the other end.


Build 14: Small Tools, Hard Rules

Two miniatures. The tools are small enough to finish in a sitting. The rules are what make them worth trusting.

The screener that demands receipts

A friend told me his dividend rules in person: the yield band he trusts, the payout streak he requires, the balance-sheet smells he walks away from. Rules he had applied by hand, stock by stock, for years. One afternoon soon after, I froze those rules into a versioned config file with a house rule written at the top: thresholds change only by versioned edit, never to make the results fit.

Then the machine read the whole market. All 2,450 Hong Kong-listed stocks in about ten minutes. The funnel: 2,450 in, 26 past the yield and streak gates, 14 past the balance sheet, 2 past everything. Two. And every one of the 2,448 eliminations carries a receipt: the filing or the raw dividend record that killed it, one row per corpse.

The receipts are the product, because headline yield numbers lie. A stock showing 35 percent yield on public screeners resolved to a true 0.0: six years of nothing, then one special payout that fooled every average. Another showing 65 percent resolved to 6.9, a per-10-share error in the raw feed. A terrifying 95 percent dividend cut turned out to be a phantom: the prior year was inflated by a one-off land sale and the ordinary dividend never broke. A 7.1 percent yield with a 15-year streak died anyway, for negative book equity funded by debt. The screen does not trust a single reported number. It rebuilds every dividend from the raw per-payment record, checks itself against a hand-built answer key (ten golden stocks, zero hard failures), and sends the ambiguous cases to a human queue instead of guessing. “I don’t know” is a real verdict here, not a rounding of yes or no.

My favorite catch was not seeded at all. One soft rule, written for one flagged stock, about cash bleeding sideways under a controlling family: it auto-caught two unrelated tickers under the same family that nobody told it to look for. Write the rule, not the answer, and the rule keeps working after you stop.

None of this is investment advice. It is bookkeeping with teeth.

The watcher that cannot click

The second miniature guards an appointment. A recurring booking on a public system where earlier slots sometimes open and vanish, the kind you catch by refreshing a page at the right second or not at all. So a watcher polls on a fixed schedule, compares every open slot against a written policy for what counts as genuinely better, and alerts only on a true match.

The design rule that matters: it is structurally incapable of booking. There is no code path from “found one” to “took it.” It never logs in, never stores a credential, and the human click on the real page stays the human’s, on purpose, not as a missing feature.

Its production numbers are a portrait of honest automation. 89 polls: 84 said nothing changed, 5 saw the same earlier slot across consecutive passes, and a cooldown collapsed those five sightings into exactly 1 alert. Out of 72 watch starts, 70 hit the session-expiry boundary and stopped to ask for a human. Read that again: the most exercised code path in production is the safety boundary. That is not a bug report. That is the design working, tested by reality dozens of times, guessing zero times. Under it all, 88 test cases against roughly 2,400 lines of watcher, near one line of test per line of code.

Silence means healthy. It speaks only when something needs me: a better slot, or a dead session that needs a human. Everything else is a quiet line in a log.

The shared rule

These two tools have nothing in common except the thing that matters: each one encodes a hard rule it will not bend. The screener may not soften a threshold to please the results. The watcher may not act without a valid session, and may not book at all. Small tools are easy to write now; the model writes them in an afternoon. The craft has moved to the rules: what the tool must never do, held so firmly that reality can lean on it 70 times out of 72 and nothing gives.

Ask my dividend friend which part he trusts. It is not the code. It is the row with his rule’s name on it, next to the stock it killed.


Build 15: My Running Coach Argues With a 1998 Textbook

A training engine that must agree with the published science before it may speak, and the race it just coached.

Why it exists

Singapore’s December race draws over 55,000 runners into air that does not forgive. Heat casualties rise every year, and no mainstream training app adjusts for the thing that actually breaks runners here: humidity. Generic plans are written for weather Singapore does not have. I am training for a sub-4 marathon in this air, so I built the coach I could not buy.

What it is

There is no app. runcoach is a coach I talk to inside Claude Code, in the same terminal I build in. The conversation is the interface, and the loop is the product:

It pulls my training history off my Garmin watch: every run, with pace, heart rate, cadence, splits. It reads that history against Jack Daniels’ Running Formula, the pace-tables textbook from 1998 that still underwrites most serious distance training. It takes my goal, a sub-4 marathon in December, and writes the plan. It pushes the workouts back to the watch as structured sessions, so on the road my wrist tells me the next step. I run. Afterward it reads the tape and we talk: what held, what drifted, what the heart rate says the pace meant. Then we build the next week from what actually happened, not what was supposed to. Around the loop again, every week.

The trust gates

What separates it from a chatbot with opinions: the engine must agree with the published science before it may speak. At VDOT 50, the code has to reproduce the book’s own numbers before any workout goes out the door. 5K in 19:56. 10K in 41:20. Half in 1:31:31. Marathon in 3:10:40. The gate demands agreement within seconds and runs every session. If my code and the textbook disagree, the workout does not render.

Then it argues with tomorrow. Before a pace reaches me, the engine prices in the dew point: at 24 degrees of dew, roughly a 7 percent adjustment. The plan you get is not the plan from the book. It is the book’s plan, corrected for the air you will actually breathe.

Five engines, each with its own gate: fitness scoring, lactate threshold, heat, workout structure, watch ingest. About 2,400 lines across 24 files. Week one, I read all six workouts back off the device to confirm they survived the trip. The claim is not “it generates workouts.” The claim is “what arrived on my wrist is what the science said, verified end to end.”

Last Saturday it coached a real race

This month the coach got its first live start: a half marathon in evening heat. The buildup was honest about me. A work sprint ate my taper, eight straight days off, and the plan’s own gate session never ran. So the system applied its own rule and took the sub-2 attempt off the table before I could argue. At 5:55 p.m. on race day it rewrote the card around what the data could still defend: a 2:05 negative-split plan, heart rate as governor, pace as ceiling. By 5:58 the workout sat on my watch, read back and verified. The card even predicted my failure mode, in writing: you feel 5:50, you run 5:58. I broke rule one in the first kilometer anyway.

At kilometer 10 my right calf began to twitch. I spent the back half managing a cramp that never fully landed, and finished in 2:14:53. Both time goals missed.

Then came the part I actually built the coach for: the autopsy. Front eleven kilometers, 5:58 per km at heart rate 171. Back ten, 7:01 at the same 171. The same engine, 63 seconds per kilometer slower. The coach decomposed the tape: cadence down 1.8 percent, stride length down 13.2 percent. That signature is not fitness and it is not heat. It is push-off. A calf protecting itself, all the way down to the 151-cadence shuffle at kilometer 16 where the heavy cramps finally landed. Two verdicts followed. The miss was written during the missing fortnight, not on the course. And the model in my head, cramps only come after 21k, was wrong in a way the data could name: cramping tracks intensity relative to trained state, not distance. My marathons never cramped early because I was trained for them. Saturday I was not, and my calves knew before I did.

The next morning it built my 19-week block for December’s marathon, then audited its own draft against Daniels and cut my goal. Sub-4 demands a fitness level I have never demonstrated; the defensible target is 4:00 to 4:05. My coach read my race, checked the textbook, and revised my dream downward, with citations. That is exactly what I built it for.

The twelve minutes

At 23:43 one July night I committed the personal version: my lactate test, my watch, my race. At 23:55 I committed a multi-user SaaS scaffold: signup, onboarding, a plan engine, coach chat, payments with a trial gate. Twelve minutes between commits. The next morning, one more session added a second AI provider behind the first with a graceful fallback behind both, so a runner mid-taper never asks a question into silence.

That was the whole arc: under 11 hours of commits from personal script to priced product. This is what the new economics feel like from the driver’s seat. The scaffold is not the achievement. The achievement is that scaffolds stopped being a reason to wait.

Unlaunched, on purpose

One clarification, because the two layers matter. The coach you just read about is not a product waiting to launch. It is a working tool I train with every day, and it needs no launch to exist. The unlaunched thing is the spinoff: that twelve-minute web-app scaffold with signup and payments. It has no public URL and no users, and it stays parked, because my constitution says no launches without a named user, and no runner has asked yet.

If you are training through Singapore heat for December: the first conversation is free and it will be about your training, not my app.

What you can steal

  1. Verify generated plans against the source science programmatically, not by eyeballing. If the code cannot reproduce the textbook’s own tables, the code is wrong, whatever the vibes say.
  2. Price the weather. Any outdoor-training product that ignores dew point is coaching a city that does not exist.
  3. Provider redundancy is a kindness: first AI answers, second picks up mid-conversation if the first dies, and a grounded fallback answers if both do. Never leave a user talking to a spinner.
  4. Build the scaffold in twelve minutes if you like. Launch it when a named human wants it. The order matters.

If one of these makes you think of your own document pile, feed, or workflow: @phuakuanyu. First reply is a real look at your problem, not a pitch.