PM Craft · Agentic Systems

Practice in Private. Publish in Public.

“Most PM judgment never gets graded. Conrad forces a prediction, a confidence number, and a real resolution date - so being wrong is data, not a vibe.”

A private tool - the product and security reasoning behind it, including what’s working and what isn’t yet.

Role
Sole builder (private tool)
Built
Claude Code as the engineering team
Timeline
2025
Private · Tailscale-only
18
published posts, 573K+ impressions
6
real decisions loaded, scored like a forecast
2
real fixes: a mechanical voice, an access-boundary gap
1
product decisions and build
01

The feedback loop problem

Most product decisions never get resolved in any way you can learn from.

No real grading
Confidence and correctness never get compared against each other - you can be consistently overconfident for years and never notice.
Feedback arrives too late
By the time an outcome is clear, the reasoning behind the original call is already gone.
Judgment and writing stayed separate
Practicing decisions privately never fed into how I communicate them in public.
02

One brain, two jobs

The same memory core feeds a private judgment practice and a public content engine - neither is a copy of the other's data.

One shared memory, two separate halves
What it’s learned
past posts · my voice, distilled · what’s worth writing about
feeds
Judgment training
private · cases, decisions, scenarios
feeds
Content engine
public · LinkedIn drafts

Neither side is a copy of the other’s data - both read from the same memory, so a pattern learned practicing a real decision privately and a pattern learned about my own writing voice both come from one place, not two drifting ones.

Exhibit: Conrad's actual home screen, live right now
Conrad's real home screen showing Today's Call for judgment training and a live content signal side by side, plus real usage stats: 0 day streak, 0 judgment reps, 10 angles queued

Both halves on one screen - Today's Call for judgment practice, a live signal feeding the content engine. The 0s here are real, not staged - see Section 03.

03

How judgment gets graded

Not a diary. A forecasting exercise, applied to real decisions.

Grade with a Brier score, not right/wrong
What I didn’t do insteadA plain decision journal
WhyWriting a decision down without a graded prediction is a diary, not a feedback loop.
Real historical cases, real documented outcomes
What I didn’t do insteadGeneric hypothetical scenarios
WhyDecision archaeology on cases with a real, checkable answer teaches calibration a made-up scenario can’t.
Bring-your-own real decisions, resolved on a set date
What I didn’t do insteadOnly pre-loaded case studies
WhyThe decisions that matter most are the ones I’m facing, not someone else’s.

Honest state of this half: the scoring pipeline is real and correct - six cases loaded, calibration math verified - but I haven’t logged a real decision through it yet. Built before it was a habit, not the other way around.

04

When the writing went quiet, then stale

The content half went about five weeks without a post. Restarting it meant first working out why the writing had gotten predictable enough to stop wanting to publish it.

The diagnosis
Every draft was being shown the account’s own highest-reach posts as “the voice” to imitate. That just kept reinforcing whichever format had gone viral once, not the most honest response to a given topic - so everything came out sounding like the same formula.
The fix
Rewrote the voice guidance around three real shapes instead of one, and changed which past posts get shown as reference so quieter, more personal posts count as much as the loud ones. Verified against a real draft that came back in a shape the old system had never produced.
Exhibit: real monthly posting trend
Conrad's real analytics dashboard: monthly post count and engagement trend from 2025 through 2026, showing the gap where the account went quiet

The actual trend line - notice the gap before the last bar. That's the five weeks this section is about.

05

How much autonomy to give the AI writer

Replacing a single prompt with an agent that can check its own facts before writing - the trust boundary I drew, not the plumbing underneath it.

What the AI writer can do
The AI writer
drafts posts on my behalf
can only
Look things up
past posts · my voice profile · content plan · goals
Can’t post on my behalf. Can’t change any data. Can’t run commands - not a rule it follows, a capability it doesn’t have.
Let it look things up, not act
What I didn’t do insteadGive it the same broad access I’d give a coding assistant
WhyWriting on my behalf, it should be able to get its facts right - not be able to change anything. A mistake stays a bad draft, never a bad action.
Make failures loud
What I didn’t do insteadQuietly swap in a different writer if this one fails
WhyA draft that looks real but wasn’t is worse than getting no draft and knowing it.
Exhibit: a real draft, mid fact-check
Conrad's real Draft Workshop showing AI-generated LinkedIn drafts tagged 'to verify' and 'grounded' - claims the agent flagged for me to check before anything gets published

This is “can look things up, not act” in practice - drafts come back with claims flagged to verify, never auto-published.

06

What I found when I checked my own work

Before rolling this out further, I had it reviewed like a production system, not a side project.

Caught before it shipped wide

The review found a real gap: one of the safety settings I believed was fully applied wasn’t - the process had more system access than I’d designed for. Low real-world risk, since it still couldn’t touch files or run anything, but a claim the whole design leaned on turned out to be only half true. Fixed before it became load-bearing for anything else. The actual lesson wasn’t about this one setting - it’s that a private, one-person project doesn’t get to skip the review a shared product would get, just because no one else is watching.

07

Why this stays private

A deliberate choice, not a gap.

Conrad runs against real business content, real credentials, and my own real decision practice - the same reasons RoleRadar and Safar are public are exactly why this one isn’t. This case study is the entire public surface of it: the real product decisions, not a demo account or a sanitized clone.

Get in Touch

Let's talk about
how you think about product decisions.

If calibrated judgment and product decisions made under real uncertainty are things you care about too, I'd love to talk.