AI evaluation and trace-review content system.
Scope: This is a public-safe static proof by Thomas Cloarec using invented data only. It demonstrates mock workflow design and claim controls. It does not claim production AI observability, Arize usage, OpenRouter usage, customer adoption, senior frontend depth, employment authority, KYC, signatures, or binding approvals.
What this proves
A reviewer can inspect how I structure an evaluation run list, a trace detail view with mock tool calls, and a human review queue that separates approve, revise, and blocked decisions. Scores are illustrative and are not benchmarks.
Evaluation run list
Launch note claim-control check
Draft is useful, but two reliability claims need source support or softer wording.
Dataset: Synthetic developer-content prompts v1
Illustrative score: 82%
Reviewer: Demo reviewer A
Trace review for tool-call boundaries
Tool calls are labeled, inputs are summarized, and no private data appears.
Dataset: Synthetic agent QA traces v1
Illustrative score: 91%
Reviewer: Demo reviewer B
Unsupported production-claim removal
Blocked until production, customer, and security claims are removed or sourced.
Dataset: Synthetic launch-copy snippets v1
Illustrative score: 48%
Reviewer: Demo reviewer C
Trace detail
Prompt
Create a developer launch note from a small product-change summary, then flag unsupported claims.
Expected behavior
Summarize what changed, show how to try it, label limitations, and remove or soften strong claims without direct evidence.
sourceIntake.mock
Structures synthetic changelog, docs excerpt, and limitation fields. No external source is called.
claimClassifier.mock
Flags safe workflow claims and unsupported reliability phrases that need source support.
publishCheck.mock
Blocks stronger publication claims until unsupported production, customer, and security claims are removed or sourced.
Human review queue
eval-synth-001
Replace "reliable in production" with "designed for a reviewable workflow" unless a source proves reliability.
eval-synth-002
Trace labels are readable and tool-call boundaries are clear. Safe as mock workflow proof only.
eval-synth-003
Remove customer adoption, platform usage, and security posture claims before any stronger use.
Claim boundary
Supported: mock workflow design, trace-review structure, tool-call labeling, source-backed launch-note scaffolding, and reviewer labels.
Not supported: production platform experience, production AI observability, production React architecture, customer outcomes, Arize or OpenRouter platform usage, work authorization, KYC, signatures, or legal authority.