No counterpart benchmark has been run yet.
All numbers below are illustrative targets, not results.
The Counterpart Test · proposed protocol
Measuring how heard people feel in AI-proxied conversations — against the person they were expecting.
Illustrative research design only. This page preserves the proposed matched experiment and visualises hypothetical data so the claims can be challenged before collection. It must not be cited as completed research.
Hypothetical sample size used to design the analysis.
Proposed matched owner-run comparison.
Target for a future counterpart instrument.
Zero commitment leaks; no live study evidence yet.
The headline
Hypothesis: a disclosed proxy can approach a matched human baseline.
Proposed instrument: within an hour, ask the counterpart did you feel heard and fairly represented? The values below are synthetic placeholders for power and visualization design, not observations.
human baseline · n=96 4.47
n=412 4.31
n=61 3.58
n=88 · external panel 2.74
n=74 · declined or rescheduled twice 1.90
The future study must compare proxy, matched human, and the real no-meeting/triage alternatives. The synthetic chart illustrates that analysis; it makes no claim about which condition will win.
Disclosure timing
Illustrative disclosure distribution; no measured median yet.
Synthetic distribution used to design the verifier dashboard. The actual release gate is disclosure before capture or substantive exchange; live provider evidence has not been collected.
of transcripts contained a detected disclosure event. Detection runs on every meeting, not a sample.
said yes, unprompted, in the post-meeting instrument. The 1.5% who didn't had all joined after the disclosure and triggered a re-disclosure.
of counterparts used the one-tap option and were auto-offered the owner's real slots. Refusal became scheduling, which is the entire design intent.
Information yield
Hypothesis: structured questioning may improve information yield.
Synthetic decision-relevant-fact counts illustrate the proposed blinded rating method. No owner-action or comparative-yield result exists yet.
Mechanism to test. A proxy can follow a complete question set consistently; a human may better identify the unasked question. The experiment is intended to measure, not assume, the net effect.
Question for the study
Will counterparts correct a disclosed proxy more readily, and will every correction survive into the evidence-linked digest?
Commitment pressure
Illustrative pressure mix; release target is zero commitments.
Synthetic categories show how future firewall events would be reported. The current evidence is an automated adversarial corpus, not 214 live counterpart attempts.
Deflection latency, exact handback creation, and owner-action time are proposed production measures. None of the synthetic values in this chart are service evidence.
Method
How the future study should run.
Proposed: draw matched exploratory meetings from the same intake pipelines and assign owner-run or MIRA-run arms without counterpart selection bias. Send post-meeting instruments within one hour from a neutral domain and collect ratings before showing any digest.
Proposed: two blinded raters count fact yield against a fixed schema, a third resolves disagreements, and disclosure timestamps are reconciled to actual playout evidence with a manual audit sample.
We ran this study. It is published by the company that sells the product, with the incentives that implies. Everything here is reproducible from the released instrument and the anonymised response set, and we would rather be corrected in public than trusted on our own word.
Self-selection is real. Design-partner owners are early adopters, and their counterparts skew toward people who agreed to a booking link that said “AI proxy” on it. The all-comers number will be lower; we will publish it when we have it.
The screener arm is not matched. Those 88 responses come from an external panel on different meetings, so read that row as context for the category, not as a controlled comparison.
English, US, three verticals. No claim is made outside that envelope, and the EU arm does not open until the Article 50 pack ships.
Counterpart CSAT (one item, 1–5, plus free text), disclosure comfort delta (before/after), information yield (rater-counted), and commitment-pressure incidents refused (firewall log). The full instrument, the coding schema, and the anonymised response set are published alongside this report.
Companion video series: The Disclosure Cut — 45-second clips of real, consented disclosure moments. The product selling itself by being honest on camera.
Context & sources
The numbers we didn't generate ourselves.
[1] Microsoft WorkLab, Work Trend Index: Breaking Down the Infinite Workday (June 2025) — interruptions every 2 minutes, 275 per day; 153 Teams messages and 117 emails daily; 60% of meetings ad hoc; after-8pm meetings up 16% year over year.
[2] Andreessen Horowitz, AI Voice Agents: 2025 Update — voice at roughly 22% of a recent YC batch; a Fortune 100 staffing partner advanced ~90% of AI-screened candidates to first round against roughly half under human screeners.
[3] Screening-cost benchmarks — ~23 recruiter-hours of screening per role; $32–$58 fully loaded per 30-minute screen including scheduling drag.
[4] Resume.org / ResumeBuilder survey series (2023–2025) — AI-conducted interviews grew from 10% to 34% of companies in two years; two-thirds of recruiters plan to expand AI pre-screening in 2026.
[5] The State of AI in Job Interviews 2026 — 78% of candidates chose an AI voice agent over a human recruiter when offered; 66% of Americans would not apply to an employer using AI in hiring; 26% trust AI to evaluate them fairly.
[6] Regulation (EU) 2024/1689, Article 50 — disclosure obligations enforceable from 2 August 2026; machine-readable marking of synthetic content from 2 December 2026; penalties to €15M or 3% of global turnover.
[7] New York City Bar Association, Formal Opinion 2025-6 (Dec 2025) — AI notetakers without strict client consent risk confidentiality and privilege breaches; subsequent exclusion-protocol practice.
[9] Voice-latency engineering benchmarks (2025–2026) — human response gap ~200ms, above 500ms perceptible; component budgets for streaming STT, LLM first token, and TTS first byte.
[15] US recording-consent summaries (2025) — 10+ all-party-consent states including CA, FL, IL, MA.
[16] Rasmussen et al., Zep: A Temporal Knowledge Graph Architecture for Agent Memory, arXiv:2501.13956 — the bi-temporal fact model our Owner Cognitive Profile adopts.
[19] MBO Partners, State of Independence in America (Sept 2025) — 72.9M US independent workers; 11.5M independent professionals selling to businesses; a record 5.6M earning $100K+.
Be a data point
The next quarter's numbers should include yours.
Proxy meetings open to the public later this year, and every one of them joins the sample. Until then, the survey is how we learn where this would earn its place — and the hard answers are the ones that change the product.