Case study
A daily news briefing for small business owners that turns the week's regulations, industry news and local policy into actions they can take. I took it from an idea to a paid product in production.
Owners of small businesses, real estate agents and accountants are affected by new rules and local decisions they rarely have time to follow. Compass Brief reads the sources they choose, then writes a short brief about what changed and why it matters to their own projects. Every brief ends in specific actions.
From there an action can become a step-by-step plan (BriefPlans), with Claude suggesting next steps. Alerts flag news that can't wait for the weekly brief. A law tracker follows state and federal bills and summarises each change in plain language. At the end of each quarter, the app asks how the user's goals went and carries that into future briefs.



I worked as the product manager and designer, directing an AI coding agent that wrote the code. I didn't write the code myself, and I don't claim deep knowledge of it. My job was deciding what to build, how it should look and behave, what "done" meant, and when it was safe to ship.
My background is QA automation: five years with Playwright and TypeScript. Testing is where I directed the agent most closely. I treated it like a very fast engineer: clear requirements for what to test and why, then verify everything it produced.
The agent explored the app and wrote a 182-scenario test plan; I set the priorities (must-have, important, polish) and the order to work through them. Every scenario now has a test, or a recorded reason it doesn't.
170 Playwright tests drive the real app and API on emulated Android and iPhone. The AI, payments, email and bill data are replaced by fakes that answer instantly and can be made to fail on purpose.
I proposed a movable clock, so rules that depend on time (a 60-second code resend, 3 briefs a day, the 14-day trial, a 60-day archive) are tested in seconds instead of waiting.
When a test failed, the first question was whether the app or the test was wrong. Both happened. Real bugs were fixed, test mistakes were corrected, and trade-offs were mine to decide.
Low enough to try without a second thought, and locked in for as long as the subscription stays active, so early users are rewarded for coming first. Raising prices later means new prices for new sign-ups only.
Sign-up goes straight to setup, with no checkout in the way. Subscribing during the trial charges from day one, and the app says so, so nobody is surprised by a charge.
At $7 a month, sending a user's whole history to Claude would cost more than they pay. Briefs are built on the server with fixed limits, and Claude gets a size-capped summary of each user instead of raw history.
When many users follow the same bill, it's fetched and summarised once. That keeps the law tracker's cost tied to the number of bills, not the number of users.
No pop-up confirmations anywhere. Deleting a brief, a project or a plan step hides it at once and offers Undo for six seconds; errors show next to what failed.
A deleted or archived project stops reaching Claude from the next request. Deleting an account cancels billing first, so nobody is charged for an account that no longer exists.
When testing found one theme's pill slightly below the contrast standard, I chose to keep that theme's original colours. When an article link died, I chose to keep the action and strike the link through rather than hide useful advice.
I didn't write the code, but running the product meant owning its setup and the technical decisions around it, working from the agent's step-by-step guidance.
Production and a staging copy, each with its own database branch, cache and payment mode, so test data never reaches real users. Every change goes to staging first.
Moved billing from test to live Stripe keys and set up the live webhook. At launch, the webhook pointed at an address that redirects, so payments weren't reaching the app; the fix was pointing it at the main domain and replaying the events.
DNS on Cloudflare, company email on Google Workspace with sender authentication (SPF, DKIM, DMARC), app email through Resend on a verified domain, and permanent redirects for the alternate domains.
Applied each schema change by hand on staging, then development, then production, in order, before the code that needed it shipped.
Asked for a movable clock so time-based rules (a 60-second resend window, daily limits, the 14-day trial) could be tested without waiting, and set the priorities for the test plan.
Defined what Claude sees after a user deletes or archives a project, and that deleting an account cancels billing before any data is removed.
Every change went to a staging copy of the app first. Automated checks ran on each one (type checking, linting, unit tests and a code-quality scan), and only after they passed and staging deployed cleanly did the change go to production.
Get the first paying customers and learn what they actually use. Then: a watch on production errors and AI cost per user, and a way for briefs to run for many users at once as the user base grows.