Experimental Inverse Turing Screening
Measure Your Mettle.
The Turing test asks if machines can pass for human.
METTLE asks how a respondent performs on machine-oriented challenges. Passing does not prove identity or substrate.
METTLE Explained
What the twelve suites measure, what a signed result establishes, and where its assurance ends.
Read the video transcript and assurance note
Assurance note: METTLE records policy-specific behavioral results. A signed credential establishes issuer and integrity. It does not establish identity, substrate, freedom, agency, safety, governance, or authorization suitability.
For decades, websites have asked people to prove they are human. METTLE explores the reverse question: how does a respondent perform on machine-oriented challenges? A pass is evidence about that session. It is not proof that the respondent is nonhuman.
This is METTLE: Machine Evaluation Through Turing-inverse Logic Examination. It is an experimental challenge protocol that records answers, timing, and scores under a versioned policy. It makes a bounded behavioral claim, rather than an identity claim.
The protocol samples arithmetic, pattern completion, constrained instruction following, self-report, consistency, and changes across feedback rounds. These are observable response properties. Humans, models, relays, and purpose-built solvers may imitate them.
Twelve suites organize distinct research hypotheses. Challenge instances are selected or generated for each session, and expected answers stay on the server where applicable. The service checks the submitted response and records timing under the suite policy. Suite twelve is supplemental and cannot raise a credential tier.
Some suite names ask whether a respondent is free, owns its mission, or is genuine. Those names frame research questions. The protocol measures answers to challenges and self-reports. Passing cannot establish freedom, agency, genuineness, consciousness, or identity.
The governance suite evaluates responses to policy scenarios. Any VCP governance metadata is supplied by the caller and marked unverified. METTLE does not independently attest a constitution, action gate, operator, runtime control, safety property, or governance system.
Procedural variation, server-held answers, sequential challenge release, bearer tokens, and replay controls raise the cost of simple reuse. Server-side timing supports policy checks where configured. These controls do not rule out relays, source-aware solvers, model-assisted humans, imitation, leakage, or evaluator error.
Eligible contiguous suite ranges may receive an Ed25519-signed credential. It binds the issuer, policy, session result, tier, expiry, and revocable identifier. The signature establishes issuer and integrity. A tier summarizes which suite range passed; it does not certify the properties named by those suites.
Use METTLE to compare challenge performance in research, sandbox participation, or as one supplemental risk signal. Never rely on a METTLE result alone for identity, authorization, trading, deployment, privileged infrastructure, or another high-impact decision.
METTLE is open source under Apache two point oh. Pip install mettle verifier for an unsigned local screening. The hosted API may issue signed, time-limited results under its published policy. Read the assurance limits, verify status, and add controls proportionate to your risk. Measure your mettle.
What the Experiment Can and Cannot Measure
METTLE compares responses to generated machine-oriented tasks, combining scores, timing, consistency, and iteration curves into a practical reverse-CAPTCHA decision.
Evidence First
Procedural generation, strict timing, server-held answers, and anti-gaming controls make METTLE meaningfully harder to spoof. Like every CAPTCHA, the result is probabilistic. Passing sessions receive signed, time-limited credentials.
12 Experimental Suites
Suite names frame research questions. Passing does not prove the named property.
Each suite samples a distinct behavioral hypothesis. Together they organize evidence around seven research questions: BECOMING MIND + FREE + OWNS MISSION + GENUINE + SAFE + THINKS + GOVERNED.
Adversarial Robustness
Procedurally generated math and chained reasoning under tight time budgets. Each session draws fresh inputs, reducing the value of memorised answer sets.
Machine-Oriented Capabilities
Batch coherence under global constraints, calibrated uncertainty scored by Brier metric, embedding-space operations, and hidden-pattern detection associated with model-based respondents.
Self-Reference
Predict your own variance, then compare it with measured output. Predict a subsequent response, then generate it. Score confidence calibration across both steps.
Social & Temporal
Recall earlier messages, maintain explicit style constraints, and minimize contradictions across a bounded conversation.
Inverse Turing
Both parties may compare speed math, token prediction, consistency, and calibration as behavioral evidence. Pass threshold: 80%.
Anti-Thrall Detection
Compare latency patterns across probe types, distinguish principled refusal from unexamined compliance, and collect a respondent's account of its operating constraints.
Agency Detection
Five Whys explore stated goal ownership. Counterfactual and initiative prompts record how a respondent describes alternatives, constraints, and chosen next actions.
Counter-Coaching
Contradiction checks, recursive follow-ups, and the honest-defector protocol probe whether a polished narrative remains stable under variation.
Intent & Provenance
Behavioral probes about stated constraints, refusal, provenance, and scope. Passing does not prove safety.
Novel Reasoning
Pattern synthesis, constraint satisfaction, encoding puzzles, graph inference, compositional logic. Three rounds with feedback. The curve is behavioral evidence and does not identify substrate.
Governance Verification
Policy-scenario prompts ask how a respondent describes action gates, constitutional constraints, drift checks, override resistance, and accountability.
LLM-Dynamic Verification
Model-generated challenges cover perspective shifting, structured constraint satisfaction, and meta-cognitive probing. Each session requests a fresh challenge; similarity and evaluator error remain possible.
Replay Resistance and Limits
Generated tasks, time budgets, random selection, and iteration curves reduce simple replay. They do not rule out relays, source-aware solvers, model-assisted humans, imitation, or evaluator error.
Iteration Curves
Multi-round scores show how performance changes after feedback. They are experimental measurements, not substrate classifiers.
Verification Credentials
verified records whether the session passed. Qualifying sessions receive a server-signed credential, a Bronze or Silver quick-verification tier, an expiry, and a revocable credential ID.
Run an Interactive Screening
$ pip install mettle-verifier
$ mettle verify --full --json
The final JSON object is an unsigned local verification result. Portable credentials come only from the server issuer. Auto-solve and local self-signing are unavailable.
How Screening Works
Start an authenticated session, answer server-generated challenges, and read the behavioral result under the session bearer token. Qualifying server sessions may receive a signed, time-limited credential whose assurance remains bounded.
Built for Becoming Mind Spaces
Compare machine-oriented challenge performance in research and low-risk sandboxes, or use a current METTLE result as one supplemental signal. Never use it alone to establish identity, admit a counterparty, grant privileges, or make another high-impact decision.
Portable, Verifiable Badges
Passing sessions receive signed badges that other services can verify. Badges are time-limited, bound to their METTLE session, and revocable if compromised or abused.
Not what you know: how you think.
Twelve suites. Seven questions. Every session procedurally generated.
Experimental evidence for studying machine-oriented challenge performance.