← Compass Brief

Case study

Compass Brief

A daily news briefing for small business owners that turns the week's regulations, industry news and local policy into actions they can take. I took it from an idea to a paid product in production.

My role
Founder and product lead
Background
QA automation engineer (Playwright, TypeScript)
How it was built
With an AI coding agent (Claude Code), which wrote the code under my direction
Timeline
March to October 2026
Status
Live, with subscription billing
Try it
compassbrief.app
Watch
Narrated product tour on YouTube

What it does

Owners of small businesses, real estate agents and accountants are affected by new rules and local decisions they rarely have time to follow. Compass Brief reads the sources they choose, then writes a short brief about what changed and why it matters to their own projects. Every brief ends in specific actions.

From there an action can become a step-by-step plan (BriefPlans), with Claude suggesting next steps. Alerts flag news that can't wait for the weekly brief. A law tracker follows state and federal bills and summarises each change in plain language. At the end of each quarter, the app asks how the user's goals went and carries that into future briefs.

The Brief tab: a headline about an IRS e-filing rule, a summary written for the user's business, and a NOW action with its source.
The brief. News written for the reader's own projects, ending in actions.
The BriefPlans board: an action turned into a plan with completed and suggested next steps.
BriefPlans. One action grown into a plan, with Claude's ideas for next steps.
The Laws tab: tracked bills with their status and a plain-language summary.
Law tracker. Bills followed in plain language, with alerts when they move.

My role, plainly

I worked as the product manager and designer, directing an AI coding agent that wrote the code. I didn't write the code myself, and I don't claim deep knowledge of it. My job was deciding what to build, how it should look and behave, what "done" meant, and when it was safe to ship.

What I did

  • Defined the product, its audience and the features
  • Set pricing, the free trial and the billing rules
  • Directed the visual design: themes, landing page, copy
  • Reviewed every change on a staging copy before release
  • Decided which test findings to fix, accept or defer
  • Set up and ran the services: hosting, database, payments, email, domains
  • Recorded the launch video and narration

What the AI agent did

  • Wrote and refactored the code
  • Proposed technical approaches for me to choose from
  • Wrote and ran the automated tests
  • Reported problems it found, with options
  • Kept the project's notes and test plan up to date

AI-augmented testing

My background is QA automation: five years with Playwright and TypeScript. Testing is where I directed the agent most closely. I treated it like a very fast engineer: clear requirements for what to test and why, then verify everything it produced.

Risk first

The agent explored the app and wrote a 182-scenario test plan; I set the priorities (must-have, important, polish) and the order to work through them. Every scenario now has a test, or a recorded reason it doesn't.

Real browsers, controlled world

170 Playwright tests drive the real app and API on emulated Android and iPhone. The AI, payments, email and bill data are replaced by fakes that answer instantly and can be made to fail on purpose.

Testing time itself

I proposed a movable clock, so rules that depend on time (a 60-second code resend, 3 briefs a day, the 14-day trial, a 60-day archive) are tested in seconds instead of waiting.

Verify, don't trust

When a test failed, the first question was whether the app or the test was wrong. Both happened. Real bugs were fixed, test mistakes were corrected, and trade-offs were mine to decide.

182scenarios, prioritized by risk
170end-to-end tests on phones
~400unit tests
15defects found and fixed before users saw them

Decisions I made

A $7 founder price, locked in

Low enough to try without a second thought, and locked in for as long as the subscription stays active, so early users are rewarded for coming first. Raising prices later means new prices for new sign-ups only.

A 14-day trial with no card

Sign-up goes straight to setup, with no checkout in the way. Subscribing during the trial charges from day one, and the app says so, so nobody is surprised by a charge.

Keep AI costs below the price

At $7 a month, sending a user's whole history to Claude would cost more than they pay. Briefs are built on the server with fixed limits, and Claude gets a size-capped summary of each user instead of raw history.

One shared copy of each bill

When many users follow the same bill, it's fetched and summarised once. That keeps the law tracker's cost tied to the number of bills, not the number of users.

Undo instead of "Are you sure?"

No pop-up confirmations anywhere. Deleting a brief, a project or a plan step hides it at once and offers Undo for six seconds; errors show next to what failed.

Respect what users remove

A deleted or archived project stops reaching Claude from the next request. Deleting an account cancels billing first, so nobody is charged for an account that no longer exists.

Know which trade-offs to accept

When testing found one theme's pill slightly below the contrast standard, I chose to keep that theme's original colours. When an article link died, I chose to keep the action and strike the link through rather than hide useful advice.

The technical side I handled

I didn't write the code, but running the product meant owning its setup and the technical decisions around it, working from the agent's step-by-step guidance.

Two environments

Production and a staging copy, each with its own database branch, cache and payment mode, so test data never reaches real users. Every change goes to staging first.

Payments in production

Moved billing from test to live Stripe keys and set up the live webhook. At launch, the webhook pointed at an address that redirects, so payments weren't reaching the app; the fix was pointing it at the main domain and replaying the events.

Domain and email

DNS on Cloudflare, company email on Google Workspace with sender authentication (SPF, DKIM, DMARC), app email through Resend on a verified domain, and permanent redirects for the alternate domains.

Database changes

Applied each schema change by hand on staging, then development, then production, in order, before the code that needed it shipped.

Testing strategy

Asked for a movable clock so time-based rules (a 60-second resend window, daily limits, the 14-day trial) could be tested without waiting, and set the priorities for the test plan.

Data and AI decisions

Defined what Claude sees after a user deletes or archives a project, and that deleting an account cancels billing before any data is removed.

How it shipped

Every change went to a staging copy of the app first. Automated checks ran on each one (type checking, linting, unit tests and a code-quality scan), and only after they passed and staging deployed cleanly did the change go to production.

89changes merged, each reviewed
4automated checks on every change
2environments: staging, then production
7months from idea to paid launch

Problems testing caught before users did

Built with

What I'd do next

Get the first paying customers and learn what they actually use. Then: a watch on production errors and AI cost per user, and a way for briefs to run for many users at once as the user base grows.