Engineering intelligence for the AI era
Your team ships
3x more code.
Is their judgment getting any better?
Fleur measures what AI cannot — the human judgment, decisions, and critical thinking behind every line of code.
Early access for engineering teams. No spam.
Team B score declined 55→38
AI direct-accept rate at 64%
Works with the tools you already use
01 — The problem
The old metrics stopped working.
PR count. Commit volume. Story points. These used to tell you how your team was doing. Then AI coding agents arrived.
One engineer can now produce a week's worth of code in a single day.
The numbers go up.
But does the value?
Volume is up, value is unclear.
PRs tripled. Did delivery truly triple, or did AI just generate more boilerplate?
You can't see who did what.
You know the team uses Cursor and Copilot. You can't see how much was AI-generated, how much was revised by humans, and how much involved real architectural decisions.
Judgment capability is invisible.
When AI writes most code, an engineer's core value becomes judgment. You have zero data on whether this capability is growing or eroding.
02 — What is Fleur
So we built something different.
A platform that measures “human judgment quality” in AI-assisted engineering.
Not how much code was written.
How much real thinking happened.
03 — The work
See what's actually happening.
Team Judgment Overview
Rubber-stamp reviews up 12%
Team B score declined 55→38
AI direct-accept rate at 64%
One glance at your team's judgment capability structure and trends.
Refactor authentication module
Judgment Score: 82/100Judgment Timeline
AI generated initial approach (400 lines)
Human: rejected approach, redesigned architecture
AI: generated code on new architecture
Human: fixed 3 security boundary issues
Review: deep review, 2 rounds
Significant human judgment. AI was a tool, not the author.
Drill into any PR. See exactly what AI did and what humans judged.
04 — The method
How Judgment Score works.
We decompose human judgment into five observable dimensions. No black box — every score traces to specific PR behavior.
How deeply humans revise AI-generated code. Direct accept = zero judgment. Major refactoring = high judgment.
Cross-module, system-level technical decisions. The most irreplaceable human capability.
Substance of human code review. Detects rubber-stamping.
Quality of problem definition. AI can solve problems. Humans must define them.
Catching hidden problems in AI-generated code. AI produces confident errors. Catching them is human.
Judgment Score is not for ranking individuals. It's for understanding team-level judgment structure, trends, and AI collaboration quality. Measurement should empower, not surveil.
05 — Is it for you
Is Fleur right for your team?
Fleur is for you if:
- —Your team is 20+ engineers using AI coding tools (Cursor, Copilot, Claude Code)
- —You're a CTO or engineering lead who feels traditional metrics are no longer enough
- —You're concerned your team may be losing independent judgment capability
- —You want to understand what your AI tool investment is actually achieving
Not quite yet if:
- —Your team is under 10 people — you can just read everyone's PRs
- —You only want a DORA dashboard — try Swarmia or LinearB
- —You want to monitor individuals for performance reviews — we don't do that
- —Your team doesn't use any AI coding tools yet
06 — Invitation
Be among the first to see
your team's judgment quality.
We're building Fleur with design partners from leading engineering teams. Join for early access and an exclusive team judgment assessment report.
We'll never share your email. Unsubscribe anytime.