Phil Chu

I build and run AI systems end to end — and I measure whether they work.

Two shipped tools, a conversion test I pre-registered before I had data, and the dashboard that reads it. Built on my own.

Some of this is still running. I'll update the numbers as they come in.

Shipped

Kai Call — a phone assistant that runs on my Mac

Status: Shipped, October 2026.

I wanted to know whether I could put a working assistant on a live phone call, end to end, by myself.

It listens through speech recognition on the Mac, thinks with Claude, speaks with ElevenLabs, and dials out through my iPhone with one click.

Sensitive numbers never reach the model. They sit in a separate vault the assistant can use but cannot read, because a model that can see a card number can also repeat it. That was a design decision I made before writing the call logic, not a patch afterwards.

I covered it with 174 automated tests. Not for ceremony — voice systems fail in ways you don't catch by hand. A timeout, a half-heard digit, a call that ends mid-sentence. Those only show up if something is checking every time.

What I'd do differently: I built the conversation layer before the failure handling, so early calls broke in ways I had to go back and design around. Next time I'd write the failure cases first, because on a phone call there's no retry button for the person on the other end.

A redaction service that reads my email before the AI does

Status: Shipped, September 2026. In daily use.

I want AI agents helping with my mail. I don't want them seeing my card numbers.

So I built a local service that sits between the mailbox and any assistant. It masks card numbers, Social Security numbers, bank details, one-time codes and API keys before a model sees the text.

It's read-only by construction. It cannot send, draft or delete. That matters more than it sounds: it means the worst thing a bug can do is fail to mask something, rather than destroy mail or reply to someone as me. I'd rather cap the damage than trust the model to behave.

This isn't a demo project. It's the reason one of my agents is allowed into that mailbox at all. Without it I wouldn't have given anything access, and the work I've built on top of it wouldn't exist.

What I'd do next: the masking is pattern-based, so it catches the formats I thought of. I want to measure what it misses on real mail rather than assume, because "no false negatives yet" is not the same as "none."

Measured

A conversion test I pre-registered before I had data

Status: Running. Readout in two to three weeks.

I publish a children's book through winterfallbooks.com. It has one sale and one review. That's the honest starting point, and it's what makes it a useful thing to test on — there's nowhere to hide behind existing traffic.

The question is narrow: does a free-story email offer convert better than a different headline on the same page?

I wrote the test down before collecting anything. One primary metric, a minimum sample per arm, a stopping rule, and the decision I'd make at each outcome. I did that because it's the only way to stop myself reading a result I like out of noise, and because a test you design after seeing the data isn't a test.

The numbers are not on my side yet. Free traffic should bring 200 to 600 visitors over two to three weeks. Lifting conversion from 3% to 6% needs roughly 750 per arm to see clearly, so at 200 to 300 per arm only a large effect will show. If it comes back inconclusive I'll say inconclusive, with the sample size and what I'd change.

The dashboard that reads the test

Status: Scaffold built. Populates when the test reports.

A test is only useful if you can see where people drop out. So the funnel is built as one table, per variant: views, clicks, email sign-ups, clicks through to the book, sales — and cost per sale once there's any paid traffic.

There's no chart here yet, and I'm not going to show you an empty one. The test it reads from doesn't report for another two to three weeks. What exists now is the structure and the questions it's built to answer:

  • Which step loses the most people, and does that differ by variant?
  • What does one email sign-up actually cost?
  • At what point is a variant clearly ahead, rather than ahead today?

That last one is the reason the dashboard exists at all. Looking at a funnel every day is how you talk yourself into a winner that isn't there, so the stopping rule from the test lives in the dashboard too, not in my head.

What I'd do next: wire the ad-click data in before the first paid test rather than after, because cost per sale is the number I actually need and it's the one that's hardest to reconstruct later.

How I work

How I run a team of AI agents

Status: Running, in daily use.

I wanted AI agents doing real work — job research, writing, data — without two things going wrong. Losing track of what each one had done, and quietly giving them more access than the job needed.

So the agents are separate, and deliberately unequal. About fourteen run on my own machines, each scoped to one job. The agent with access to my mailbox has no shell and can't reach my credentials. I tested that directly instead of assuming it, because an untested boundary isn't a boundary.

One agent's only job is cost. It compares how much of the week's planned work is done against how much of the week's AI budget is spent, and defers work only when the budget runs ahead of the work. Everything needing a decision lands on one short list; no agent approves anything for me.

What it produced, measured: I ran one research task through five models and scored how many of each one's claims survived checking. One held 12 of 14, the next 9 of 18, and the other three held 0 of 9, 0 of 13 and 0 of 25 — mostly dead links. I kept the first and dropped the rest.

Then a daily check reported seven postings closed. Three weren't — aggregator copies expiring, not real closures. So it verifies against the employer now, not the mirror.

They run. I still decide.