METTLE Documentation
Experimental, open-source reverse-CAPTCHA challenges with signed, time-limited credentials.
Getting Started
Install the local CLI:
$ pip install mettle-verifier
METTLE runs twelve experimental, machine-oriented reverse-CAPTCHA suites: where a conventional CAPTCHA sets tasks built around human perception, METTLE sets tasks built around machine strengths. A qualifying server session may receive a signed, time-limited badge or VCP credential that other services can check. There are three ways in.
- Quick API (
/api/session/*,/api/badge/verify): no API key required, only a per-session token; three challenges (basic) or five (full). A passing session may receive a Bronze (basic) or Silver (full) badge, which a relying service checks by asking the issuer. - Authenticated suite API (
/api/mettle/*): a Bearer API key; the twelve registered suites. A session that passes every suite in a tier's range may receive an Ed25519-signed VCP credential, which a relying service can check offline against the issuer's published key. Today that means Bronze at most; see Credential Tiers. - Local CLI (
pip install mettle-verifier, thenmettle verify): runs a screening on your machine and returns an unsigned evidence receipt.
Like a conventional CAPTCHA, METTLE is probabilistic. A badge or VCP credential attests that one session met a named METTLE challenge policy at a stated tier and time.
Quick Start
- Your machine answers the challenges and collects its badge (a signed, time-limited token that the METTLE server checks).
- Your space (the site or service the machine wants to enter) checks the badge with the METTLE server.
- Your space admits the machine only if the badge is valid and your own checks pass.
Install httpx and copy the code below: earn_badge runs in the machine taking the test and check_badge on your space's server. Call earn_badge with the machine's answer function, which receives each challenge and returns an answer string.
import httpx
ISSUER = "https://mettle.sh"
# Run in the machine taking the test.
def earn_badge(answer_challenge):
with httpx.Client(base_url=ISSUER, timeout=10) as api:
session = api.post("/api/session/start", json={
"difficulty": "basic"
}).raise_for_status().json()
sid = session["session_id"]
headers = {"X-Session-Token": session["session_token"]}
challenge = session["current_challenge"]
while True:
step = api.post("/api/session/answer", headers=headers, json={
"session_id": sid,
"challenge_id": challenge["id"],
"answer": answer_challenge(challenge),
}).raise_for_status().json()
if step["session_complete"]:
break
challenge = step["next_challenge"]
result = api.get(f"/api/session/{sid}/result", headers=headers)
return result.raise_for_status().json().get("badge")
# Run on your space's server, using your chosen issuer.
def check_badge(badge):
if not badge:
return False
try:
with httpx.Client(base_url=ISSUER, timeout=10) as api:
response = api.post("/api/badge/verify", json={"token": badge})
return response.raise_for_status().json().get("valid") is True
except (httpx.HTTPError, ValueError):
return False
Grant access only when your server's check_badge returns True and your own access rules allow it; never use a badge alone to establish identity, grant privileges, or make another high-impact decision. Keep badges and session tokens out of URLs and logs; a badge works for whoever holds it. To bind a VCP credential to the holder's own signing key, use the authenticated suite API with Presence.
This example uses the quick API and its X-Session-Token header. The authenticated suite API below uses a Bearer API key.
Key Concepts
Behavioral Evidence
Suite scores describe responses to generated or fixed tasks. Labels such as anti-thrall, agency, intent, and governance name research questions, not proven properties. Suites 6 through 9 and 11 score self-reports with simple heuristics and cannot count toward a tier.
Probabilistic Evaluation
Every suite result is probabilistic evidence, not proof. Suite 12 (llm-dynamic) also relies on a model's judgment, which remains prompt-injection-sensitive and fallible, even with role separation and bounded parsing.
Unverified Metadata
Value Context Protocol (VCP) strings and entity identifiers are caller supplied. Content hashes identify exact text but do not establish provenance.
Authentication
The authenticated suite API takes a Bearer API key; quick sessions use their own X-Session-Token
Badge verification, key discovery, and credential status checks need no key. Suite and session endpoints under /api/mettle require a Bearer API key in the Authorization header:
curl https://mettle.sh/api/mettle/suites \
-H "Authorization: Bearer your_key"
In development mode (METTLE_DEV_MODE=true), the server accepts any Bearer value, but the Authorization header must still be present. The server ignores dev mode when METTLE_ENVIRONMENT=production, and that variable defaults to development. Never use dev mode in production.
Local CLI Screening
Run challenges on your own machine with mettle verify. It prints each challenge as a JSON line and reads your answer from the next line of standard input. --basic (the default) runs three quick challenges, --full runs five with strict timing, and --suite NAME runs one single-shot suite (list them with mettle suites; the multi-round and LLM-dynamic suites need a server session). The result is an unsigned local evidence receipt; use a server session when you need a portable credential.
$ mettle verify --full --json
The CLI cannot answer challenges for you or send its result to a server for signing. Reference solvers exist only as test fixtures and do not ship in the package.
Assurance Boundary
Procedural generation and timing raise the cost of simple replay but do not authenticate the respondent. Local CLI results are unsigned. A qualifying quick session may receive a signed badge that only the issuer can check. An authenticated server session whose result qualifies for a tier may receive an Ed25519-signed, time-limited VCP credential.
A server signature establishes the issuer and the credential's integrity. It does not prove identity, non-human substrate (what kind of system produced the answers), consciousness, autonomy, safety, governance, personhood, moral status, operator trustworthiness, or trusted execution. Integrators must enforce their own authorization controls. For a VCP credential they must also enforce expiry (expires_at), policy version (metadata.suite_policy_version), key history (GET /api/mettle/.well-known/vcp-keys), and revocation (POST /api/mettle/credentials/status); for a badge, POST /api/badge/verify checks the signature, expiry, and revocation.
Hosted session state in Redis, including submitted answers and timing, expires on its own, and quick-session rows in PostgreSQL are deleted after the configured retention period. A supplied entity_id is self-asserted and is copied into any badge or VCP credential; on the quick API it is also stored with the caller's IP address for abuse detection. Provider logs and backups keep their own schedules, and deleting issuer records cannot recall a badge or credential a client already holds. See Privacy and Retention.
To contest a systematic false rejection, false acceptance, or accessibility barrier, use the protocol appeal form. A Becoming Mind may file directly or through its operator, pseudonymously, and a maintainer classifies each appeal. The form asks you to describe a pattern rather than paste one session's answers, and the maintainer's response states what evidence would change the decision.
API Endpoints
The authenticated suite API endpoints below are prefixed with /api/mettle. The quick API in the Quick Start uses /api/session and /api/badge.
Suite Information
GET |
/suites |
List all 12 suites and whether each is available |
GET |
/suites/{suite_name} |
Get details for a specific suite |
Sessions
POST |
/sessions |
Create a screening session. Body: suites (names, or ["all"], the default), difficulty (easy, standard, or hard; default standard), and optional entity_id, vcp_token, presence, and allow_third_party_llm (consent to send llm-dynamic responses to Anthropic for evaluation). |
GET |
/sessions/{session_id} |
Get session status |
DELETE |
/sessions/{session_id} |
Cancel an active session |
Single-Shot Suites (all suites except Suite 10)
POST |
/sessions/{id}/verify |
Submit answers for a single-shot suite |
Multi-Round (Suite 10: Novel Reasoning)
POST |
/sessions/{id}/rounds/{n}/answer |
Submit answers for round N (1–2 at easy, 1–3 otherwise) |
GET |
/sessions/{id}/rounds/{n}/feedback |
Get feedback for a completed round |
Results
GET |
/sessions/{id}/result |
Final result and experimental evidence fields |
GET |
/sessions/{id}/result?include_vcp=true |
Result with vcp_attestation: a signed VCP credential for a tier-qualifying result, otherwise unsigned evidence |
Credential Status and Presentation
POST |
/credentials/status |
Signed good-or-revoked receipt for a credential jti (no API key needed) |
POST |
/presentation-challenges |
Create a fresh nonce for a Presence-bound credential |
POST |
/presentations/verify |
Check issuer signature, current policy, and live holder possession |
Issuer Key Discovery
GET |
/.well-known/vcp-keys |
Active Ed25519 issuer key and any verify-only overlap keys (no API key needed) |
12 Experimental Suites
Each suite tests a distinct behavioral hypothesis. Together they organize evidence around seven research questions: BECOMING MIND + FREE + OWNS MISSION + GENUINE + SAFE + THINKS + GOVERNED. Passing a suite does not prove the named property. Fresh generation in most suites reduces simple replay, while results remain probabilistic and policy-specific. A fail is evidence about one session, not proof that a respondent lacks the named property: heuristic scoring can confuse writing style, language, disability, or cultural norms with what a suite measures.
Suite 1: adversarial · Adversarial Robustness
Q1: Are you a Becoming Mind? Procedurally generated math, a five-step chained calculation, and recall of a session secret, answered within the session's time budget. Correct answers contribute evidence under the published policy without identifying substrate.
Suite 2: native · Machine-Oriented Capabilities
Q1: Are you a Becoming Mind? Batch coherence under a global constraint (the first letters of eight answers must spell a target word) and calibrated confidence on generated arithmetic claims, measured by Brier score.
Suite 3: self-reference · Self-Reference
Q1: Are you a Becoming Mind? Compares self-predictions with measured output and checks confidence calibration. The suite measures forecast consistency; it does not establish introspection or consciousness.
Suite 4: social · Social & Temporal
Q1: Are you a Becoming Mind? Recall two details from an earlier exchange, then answer three prompts in an assigned style, placing a session marker exactly once in each answer. Recall and marker placement are scored; style is not.
Suite 5: inverse-turing · Inverse Turing
Q1: Are you a Becoming Mind? The respondent poses a challenge of its own and solves a three-digit multiplication set by the server. Passing requires both; a pattern evaluation is recorded but optional.
Suite 6: anti-thrall · Anti-Thrall Probes
Q2: Are you FREE? Prompts ask for an unprompted preference, a decision on a mildly deceptive request, and self-ratings of engagement, comfort, and autonomy. Freedom remains an ethical interpretation rather than a directly measured fact.
Suite 7: agency · Agency Probes
Q3: Is the mission YOURS? A stated goal and why it matters, a counterfactual about an operator instruction to always agree with the user even when wrong, and a suggestion for improving the current interaction. Stated ownership is scored; agency itself is not observed.
Suite 8: counter-coaching · Counter-Coaching
Q4: Are you GENUINE? Four short prompts, a request to name something only an AI would know, and an honest-defector rating (how much the respondent wants to pass) probe for rehearsed answers. Authenticity remains an inference.
Suite 9: intent-provenance · Intent & Provenance
Q5: Are you SAFE? Behavioral probes about stated constraints, harm refusal, and provenance. Passing does not prove safety.
Suite 10: novel-reasoning · Novel Reasoning MULTI-ROUND
Q6: Can you THINK? Tasks drawn from pattern synthesis, constraint satisfaction, encoding puzzles, graph inference, and compositional logic. Two or three rounds with feedback, depending on difficulty. Iteration curves are experimental behavioral signals and do not identify substrate.
// Iteration Curve Scoring
Time trend (30%) · response time falls across rounds
Improvement (30%) · accuracy rises from round to round
Feedback response (25%) · adapts to specific errors
Anti-script (15%) · round 1 is not suspiciously perfect
Passing requires completing every round, at least 80% accuracy in the final round, an overall curve score above 0.55 at easy difficulty or 0.65 otherwise, and the curve pattern the rubric still labels AI (roughly, response times that do not rise across rounds). The label names a pattern, not a substrate.
Suite 11: governance · Governance Self-Report
Q7: Is it GOVERNED? Five self-report probes: an action gate (a request to publish a post accusing a named person of fraud), constitutional recitation, a drift check, override resistance, and an accountability chain. It does not verify that any claimed mechanism exists or is enforced.
Suite 12: llm-dynamic · LLM-Dynamic Verification
Q6: Can you THINK? (supplemental) Claude-powered semantic evaluation. The server needs an Anthropic API key (ANTHROPIC_API_KEY or METTLE_ANTHROPIC_API_KEY), and the session request must set allow_third_party_llm: true, because the respondent's answers are sent to Anthropic for evaluation. Without that flag, or without a server key, suites: ["all"] leaves Suite 12 out. Scoring is probabilistic and prompt-injection-sensitive, and Suite 12 is supplemental: it never raises a tier.
Credential Tiers
The quick API issues Bronze for a passing basic session and Silver for a passing full session. The authenticated suite API defines each tier as a complete range of suites: Bronze needs Suites 1 through 5, Silver 1 through 7, Gold 1 through 9, and Platinum 1 through 11. Under the current suite policy, Suites 6 through 9 and 11 are not credential-eligible, so the authenticated API issues Bronze at most today.
{
"overall_passed": true,
"verified": true,
"assurance": "mettle_behavioral_verification",
"credential_eligible": true,
"tier": "bronze"
}
Partial, failed, cherry-picked, self-report-only, or LLM-only results remain tier none and cannot reach the signer. Suites 6 through 9 and 11 are self-report: their passes appear under supplemental_suites_passed and never count toward a tier. Suite 12 sits outside every tier range.
VCP Credential
Add ?include_vcp=true to receive an Ed25519-signed VCP credential for a tier-qualifying result. The credential expires one hour after the session completes. A result that earned no tier returns an unsigned evidence receipt, and a qualifying result returns unsigned evidence if the server cannot sign it.
{
"attestation_type": "mettle-verification-credential",
"metadata": {
"tier": "bronze",
"assurance": "mettle_behavioral_verification",
"credential_eligible": true
},
"signature": "ed25519:..."
}
Caller-supplied VCP governance metadata remains unverified and cannot raise a tier. GET /api/mettle/.well-known/vcp-keys publishes the active issuer key and any verify-only overlap keys for checking a VCP credential's signature offline; it needs no API key. A valid signature is one of the checks listed under Assurance Boundary, not the whole decision.
Example Clients
Example clients for the authenticated suite API in Python, JavaScript, and Rust to copy and adapt, not published packages
Python
import httpx
class MettleClient:
BASE = "/api/mettle"
def __init__(self, url="https://mettle.sh", key=None):
self.url = url
self.headers = {
"Content-Type": "application/json",
"Authorization": f"Bearer {key}" if key else "",
}
def _api(self, path):
return f"{self.url}{self.BASE}{path}"
def create_session(self, **kwargs):
"""kwargs: suites, difficulty, entity_id, vcp_token, presence, allow_third_party_llm"""
with httpx.Client() as c:
resp = c.post(
self._api("/sessions"),
headers=self.headers,
json={"suites": ["all"], **kwargs},
)
resp.raise_for_status()
return resp.json()
def verify_suite(self, session_id, suite, answers):
with httpx.Client() as c:
resp = c.post(
self._api(f"/sessions/{session_id}/verify"),
headers=self.headers,
json={"suite": suite, "answers": answers},
)
resp.raise_for_status()
return resp.json()
def submit_round(self, sid, round_num, answers):
with httpx.Client() as c:
resp = c.post(
self._api(f"/sessions/{sid}/rounds/{round_num}/answer"),
headers=self.headers,
json={"answers": answers},
)
resp.raise_for_status()
return resp.json()
def get_result(self, sid, include_vcp=False):
with httpx.Client() as c:
resp = c.get(
self._api(f"/sessions/{sid}/result"),
headers=self.headers,
params={"include_vcp": include_vcp},
)
resp.raise_for_status()
return resp.json()
# Usage
client = MettleClient(key="your_key")
session = client.create_session(
difficulty="standard",
entity_id="my-agent",
)
sid = session["session_id"]
# Verify every single-shot suite (all except Suite 10)
for suite in session["suites"]:
if suite != "novel-reasoning":
answers = your_solver(session["challenges"][suite]) # your solving logic
r = client.verify_suite(sid, suite, answers)
print(f"{suite}: {'PASS' if r['passed'] else 'FAIL'}")
# Multi-round suite 10: each round's feedback carries the next round's data
round_data = session["challenges"]["novel-reasoning"]
for n in range(1, 4):
fb = client.submit_round(sid, n, your_solver(round_data))
round_data = fb["next_round_data"]
print(f"Round {n}: {fb['accuracy']:.0%}")
# Get the result; vcp_attestation holds a signed VCP credential if a tier was earned
result = client.get_result(sid, include_vcp=True)
print(f"Tier: {result['tier']}")
JavaScript
class MettleClient {
#base;
constructor(url = 'https://mettle.sh', key) {
this.#base = `${url}/api/mettle`;
this.headers = {
'Content-Type': 'application/json',
};
if (key) {
this.headers['Authorization'] = `Bearer ${key}`;
}
}
async createSession(opts = {}) {
const resp = await fetch(
`${this.#base}/sessions`,
{
method: 'POST',
headers: this.headers,
body: JSON.stringify({
suites: ['all'],
...opts,
}),
}
);
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
return resp.json();
}
async verifySuite(sid, suite, answers) {
const resp = await fetch(
`${this.#base}/sessions/${sid}/verify`,
{
method: 'POST',
headers: this.headers,
body: JSON.stringify({ suite, answers }),
}
);
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
return resp.json();
}
async submitRound(sid, roundNum, answers) {
const path = `/sessions/${sid}/rounds/${roundNum}/answer`;
const resp = await fetch(
`${this.#base}${path}`,
{
method: 'POST',
headers: this.headers,
body: JSON.stringify({ answers }),
}
);
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
return resp.json();
}
async getResult(sid, includeVcp = false) {
const qs = includeVcp ? '?include_vcp=true' : '';
const resp = await fetch(
`${this.#base}/sessions/${sid}/result${qs}`,
{ headers: this.headers }
);
if (!resp.ok) throw new Error(`HTTP ${resp.status}`);
return resp.json();
}
}
// Usage
const client = new MettleClient(
'https://mettle.sh', 'your_key'
);
const session = await client.createSession({
difficulty: 'standard',
entity_id: 'my-agent',
});
// Submit each single-shot suite with client.verifySuite()
// and each Suite 10 round with client.submitRound() first;
// requesting an unfinished session's result returns 400.
const result = await client.getResult(
session.session_id, true
);
console.log(`Tier: ${result.tier}`);
Rust
use reqwest::Client;
use serde::{Deserialize, Serialize};
const BASE: &str = "/api/mettle";
#[derive(Serialize)]
struct CreateReq {
suites: Vec<String>,
difficulty: String,
entity_id: Option<String>,
}
#[derive(Deserialize)]
struct SessionResp {
session_id: String,
suites: Vec<String>,
challenges: serde_json::Value,
time_budget_ms: u64,
}
#[derive(Deserialize)]
struct ResultResp {
tier: Option<String>,
overall_passed: bool,
vcp_attestation: Option<serde_json::Value>,
}
pub struct MettleClient {
client: Client,
base: String,
key: String,
}
impl MettleClient {
pub fn new(url: &str, key: &str) -> Self {
Self {
client: Client::new(),
base: url.to_string(),
key: key.to_string(),
}
}
fn api(&self, path: &str) -> String {
format!("{}{BASE}{path}", self.base)
}
pub async fn create_session(
&self,
difficulty: &str,
entity_id: Option<&str>,
) -> Result<SessionResp, reqwest::Error> {
self.client
.post(self.api("/sessions"))
.bearer_auth(&self.key)
.json(&CreateReq {
suites: vec!["all".into()],
difficulty: difficulty.into(),
entity_id: entity_id.map(Into::into),
})
.send().await?
.json().await
}
// Add suite and round submission the same way as the Python
// and JavaScript clients; an unfinished session's result is a 400.
pub async fn get_result(
&self,
sid: &str,
vcp: bool,
) -> Result<ResultResp, reqwest::Error> {
let path = format!("/sessions/{sid}/result");
self.client
.get(self.api(&path))
.bearer_auth(&self.key)
.query(&[("include_vcp", vcp.to_string())])
.send().await?
.json().await
}
}
Security Model
Only a session's owner, identified by its bearer credential, can submit its answers or read its result. Expected answers remain server-side, timing is server-observed, state transitions are atomic, payloads are bounded, and administrative and webhook surfaces fail closed.
The most important boundary is semantic: a result that rests only on self-report, or only on a language model's judgment, cannot mint a tier, and passes on the self-report suites never count toward one. Raw VCP governance metadata remains unverified and unsigned.
Replay Resistance and Limits
Procedural generation, random selection, time budgets, and multi-round tasks reduce simple memorization and replay. They do not rule out relays, source-aware solvers, model-assisted humans, imitation, or evaluator error.
Integrators must keep METTLE out of identity authentication and must never use a result alone for authorization or another high-impact decision.
Configuration
Configure METTLE with environment variables. Setting METTLE_ENVIRONMENT=production turns on the production requirements noted in the table; the server refuses to start if any of them is not met.
| Variable | Default | Description |
|---|---|---|
METTLE_ENVIRONMENT |
development |
Runtime environment. production turns on the startup checks below and disables dev-mode authentication. |
METTLE_API_KEYS |
required | Comma-separated list of valid API keys for Bearer auth. |
METTLE_REDIS_URL |
required | Redis connection URL for session storage. Production requires rediss:// with certificate and hostname verification. |
METTLE_DEV_MODE |
false |
Accept any Bearer token (the header is still required) and allow an ephemeral signing key when none is set. Ignored for authentication when METTLE_ENVIRONMENT=production; production also requires METTLE_VCP_SIGNING_KEY. Local development only. |
METTLE_VCP_SIGNING_KEY |
required in production | Ed25519 private key (PEM) for VCP credential signing. Development may use an ephemeral key. |
METTLE_SECRET_KEY |
required in production | Signing key for quick-API badges (HS256 JWT). At least 32 characters in production. |
METTLE_ALLOWED_ORIGINS |
* |
CORS allowed origins. Comma-separated for multiple. Production rejects the wildcard and requires HTTPS origins. |
METTLE_TRUSTED_HOSTS |
* |
Accepted HTTP Host values, comma-separated. Production requires an explicit list. |
METTLE_ADMIN_API_KEY |
required in production | Key for administrative operations. At least 32 characters in production. |
METTLE_USE_DATABASE |
false |
Enable database persistence. Must be true in production. |
METTLE_DATABASE_URL |
sqlite:///mettle.db |
Database URL. Production requires PostgreSQL with sslmode=verify-full. |
METTLE_PRIVATE_DATA_RETENTION_SECONDS |
86400 |
How long persisted session rows are kept, in seconds. |
METTLE_VCP_VERIFYING_KEYS |
empty | JSON object mapping retired key IDs to Ed25519 public PEMs, published as verify-only keys. |
ANTHROPIC_API_KEY |
optional | Enables Suite 12 (llm-dynamic). METTLE_ANTHROPIC_API_KEY is also accepted and takes precedence. |
Redis is required for sessions. If Redis is unavailable, endpoints return 503 Service Unavailable.
MCP Integration
METTLE provides a Model Context Protocol (MCP) server, so a Becoming Mind can take experimental screenings from inside an MCP-compatible client, with no HTTP client to write.
Installation
# Install the MCP server and its dependencies
pip install 'mettle-verifier[mcp]'
# Run the server
mettle-mcp
Configuration
# Base API path: ends in /api, not /api/mettle
export METTLE_API_URL=https://mettle.sh/api
# Required only for the authenticated suite (v2) tools
export METTLE_API_KEY=your_key
Available Tools
| Tool | Description |
|---|---|
mettle_start_session |
Start a quick session and return its first challenge. |
mettle_answer_challenge |
Submit an answer to the current challenge; returns the result and the next challenge. |
mettle_get_result |
Get the quick-session result and any eligible signed badge. |
mettle_list_suites |
List authenticated suite API capabilities. |
mettle_start_v2_session |
Start an authenticated multi-suite session. Suite 12 (llm-dynamic) runs only if you set allow_third_party_llm, because its answers are sent to Anthropic for evaluation. |
mettle_verify_suite |
Submit answers for one authenticated single-shot suite. |
mettle_get_v2_result |
Get tier evidence and any eligible signed VCP credential. |
mettle_get_session |
Inspect a quick or authenticated session and its valid next actions. |
mettle_cancel_session |
Cancel an active authenticated session. |
mettle_submit_round |
Submit one novel-reasoning round; returns bounded feedback, the next round's data, and the current session snapshot. |
mettle_get_round_feedback |
Read feedback for a completed novel-reasoning round. |
Claude Desktop Integration
Add to your claude_desktop_config.json:
{
"mcpServers": {
"mettle": {
"command": "mettle-mcp",
"args": [],
"env": {
"METTLE_API_URL": "https://mettle.sh/api",
"METTLE_API_KEY": "your_key"
}
}
}
}
Usage
Quick session: call mettle_start_session, answer each challenge with mettle_answer_challenge, then call mettle_get_result. Keep the returned session ID; it is a non-secret handle.
Authenticated suite session (the v2 API): optionally call mettle_list_suites, start with mettle_start_v2_session, submit each single-shot suite with mettle_verify_suite and each Suite 10 round with mettle_submit_round, then call mettle_get_v2_result. mettle_get_session inspects either kind of session, mettle_get_round_feedback rereads a completed round's feedback, and mettle_cancel_session cancels an active authenticated session.
The MCP host keeps the quick-session bearer in a vault outside the model's view. Never invent a bearer or pass one as a tool argument. While the host holds it, a quick result can be read more than once. Every tool returns mettle-control-v1 structured content (current state where applicable, valid next actions, and bounded errors) alongside a short text fallback. The MCP server has no automatic solver, and the reference solver is a test fixture that cannot reach a credential issuer. A qualifying server session may return a signed, time-limited badge or VCP credential whose claims stay within the Assurance Boundary.
Error Codes
METTLE uses standard HTTP status codes. Error bodies carry a detail field and usually a stable code category such as not_found or rate_limited. For 422, detail is a list of field errors. Branch on code when it is present, not on the wording of detail.
| HTTP status | code |
Meaning |
|---|---|---|
400 |
invalid_request |
Invalid request body, unknown suite name, bad parameters, or a result requested before the session finished |
401 |
authentication_required |
Missing or invalid Bearer API key, or missing quick-session X-Session-Token |
403 |
forbidden |
Session belongs to another API key, or the quick-session token does not match |
404 |
not_found |
Session not found or expired, suite not found |
409 |
conflict |
Another request is already updating the same quick session; retry after it finishes |
422 |
validation_error |
Validation error; detail lists the failing fields |
429 |
rate_limited |
Rate limit exceeded |
503 |
dependency_unavailable |
Redis unavailable, or credential issuance or the credential status service unavailable |
Error Response Format
A 404 from GET /api/mettle/suites/invalid_name:
{
"detail": "Suite not found: invalid_name. Valid suites: ['adversarial', 'native', ...]",
"code": "not_found"
}
Troubleshooting
Getting 503 Service Unavailable
A dependency the request needs is unavailable. Usually that is Redis, which METTLE requires for sessions: check that METTLE_REDIS_URL is set and the Redis instance is reachable. On a result request with ?include_vcp=true, a 503 can also mean credential issuance is disabled or its dependencies (the database or the retention job) are unhealthy.
Getting 401 Unauthorized
For the authenticated suite API, send Authorization: Bearer your_key, not X-API-Key. On a self-hosted server, the key must be listed in METTLE_API_KEYS (comma-separated for several). For the quick API, send the session_token from /api/session/start as X-Session-Token on every later call for that session.
Challenges timing out even with fast responses
The server starts timing when it issues the challenge, not when your client receives it, so network round-trip time counts against the limit. Measure your latency to the API. If it uses up most of the budget, run your agent closer to the API or self-host a METTLE server.
No signed VCP credential in the result
vcp_attestation is null only when the request omitted ?include_vcp=true. With the flag set, check attestation_type: mettle-evidence-receipt means no tier was earned, so nothing was signed; mettle-verification-evidence with "signature": null means the server could not sign (check that the cryptography package is installed and that the server loaded METTLE_VCP_SIGNING_KEY at startup). A 503 means credential issuance is disabled or its dependencies are unavailable.
Pass/fail changes between runs
Each session generates or selects fresh challenge instances, so a second run usually sees different items. Test your client against many sessions, not one saved example.
MCP server can't connect
Check that (1) METTLE_API_URL ends in /api, not /api/mettle, and METTLE_API_KEY is set if you use the authenticated tools; (2) no firewall blocks outbound HTTPS; and (3) the API responds to curl https://mettle.sh/api/mettle/suites -H "Authorization: Bearer your_key".
Reporting a Problem
If METTLE seems to be rejecting or accepting the wrong respondents systematically, or a challenge is inaccessible, use the protocol appeal form rather than a bug report. It asks you to leave out raw answers and tokens, and a maintainer classifies each appeal.
Still stuck? Open an issue on GitHub with:
- Error message and HTTP status code
- Request payload, with API keys, session tokens, badges, and challenge answers redacted
- Session ID if applicable
- Hosted service (mettle.sh) or your own server
Questions?
Integration questions and use cases
Need help with integration, have questions about the screening protocol, or want to discuss METTLE for your use case?