Standard · v1.0

Recipient Intelligence Index

A measurement framework for Recipient Intelligence in AI communication systems. Scoring 0–10 across five dimensions. Updated quarterly.

Version
1.0
Published
June 19, 2026
Last updated
June 19, 2026
Next update
September 17, 2026

What the RII is

The Recipient Intelligence Index is a measurement standard for evaluating whether an AI communication system has genuine Recipient Intelligence — or only the appearance of it. It exists because the category needs a benchmark: a fixed, citable way to compare systems beyond marketing claims. This document is a living standard, versioned and updated on a quarterly cycle.

The five scoring dimensions

Each dimension is scored from 0 to 10, for a maximum of 50. A high total means a system models the recipient and acts on that model; a low total means it writes well but writes blind — the signature of Recipient Blindness.

1. Recipient State Persistence
Does the system maintain durable context on the recipient across sessions? A 0 starts every message from scratch; a 10 carries a per-person model — archetype, preferences, history — between every interaction.
2. Adaptation Depth
Does it change message structure, not just tone? A 0 swaps a few words; a 10 reorders the message, changes what leads and what's cut, and matches length and directness to the recipient.
3. Feedback Integration
Does it update its recipient model based on response behavior? A 0 never learns; a 10 treats replies, edits, and silence as signal and revises the model over time.
4. Archetype Differentiation
Does it distinguish between recipient types, or apply a single adaptation? A 0 treats everyone the same; a 10 maps recipients to distinct communication archetypes with separate guidance.
5. Relationship Health Tracking
Does it detect stalls, drift, and reconnect opportunities? A 0 sees one message at a time; a 10 tracks the relationship across the thread and surfaces what it needs next.

Methodology

The scoring procedure is documented below. See Qualia's methodology page for the broader scientific foundations, public precedents, and boundaries that inform this work.

1. Definitions

Recipient Intelligence: the degree to which an AI communication system constructs and uses a working model of the message's recipient — as opposed to optimizing exclusively for the sender's voice, intent, or stated goal.

The RII operationalizes this into five measurable dimensions, scored independently, each worth up to ten points:

1. Recipient State Persistence (0–10)
Does the system retain information about a specific recipient across sessions, or does each message start from zero? Scored from no persistence (0–2), to persistence requiring manual re-entry (4–6), to automatic persistence and update across sessions (8–10).
2. Adaptation Depth (0–10)
Does the system restructure message content, tone, and framing for the recipient, or only substitute surface-level word choices? Scored from lexical substitution only (0–2) to full structural and rhetorical adaptation (8–10).
3. Feedback Integration (0–10)
Does the system treat recipient replies, edits, and non-response as signal that updates its model of the recipient? Scored from no feedback loop (0–2) to automatic model revision from implicit and explicit signal (8–10).
4. Archetype Differentiation (0–10)
Does the system distinguish between meaningfully different recipient types with separate guidance, or apply one generalized profile? Scored from a single undifferentiated default (0–2) to multiple distinct, evidence-based archetypes (8–10).
5. Relationship Health Tracking (0–10)
Does the system model the relationship a message belongs to — history, trajectory, current state? Scored from no relationship-level model (0–2) to a persistent, longitudinal relationship state (8–10).

Maximum score: 50.

Interpolation logic. The rubric anchors 0, 5, and 10 carry written qualitative definitions. Intermediate scores follow a consistent interpolation logic applied uniformly across all five dimensions: 1–2 indicate a trace of the capability (partially documented, not operational in standard use); 3–4 indicate the capability is present but manual, shallow, or conditionally available; 6–7 indicate structural implementation with one notable documented gap; 8–9 indicate full or near-full implementation where the remaining gap is in scope or completeness, not presence. Applied example: Gong's Adaptation Depth is a 3 because AI Composer grounds drafts in conversation history (above generic rewriting, which would be 0–2) but does not structurally reorder messages based on recipient cognitive style (which would be required for the 5-point anchor).

2. Scoring Procedure

Each tool was evaluated through direct product use (sandbox accounts, public demos, or documented API behavior where hands-on access wasn't available), cross-referenced against public documentation and published case studies. Scores reflect default/standard configuration — the question is what a typical user actually gets, not what is theoretically possible with custom engineering.

3. A Note on Independence

Version 1.0 scoring was performed by Qualia's founding team, including scoring our own product. We did not use an independent third-party scorer — that's a real limitation, not a hidden one. We've invited methodology review from outside researchers, and we're publishing the full rubric so anyone can re-score any tool, including ours, and tell us where they land differently.

4. Versioning

RII 1.0 published June 19, 2026. Revisions ship quarterly; the next is scheduled September 17, 2026. Each version is permanently archived at a versioned URL — scores are never silently updated. Material changes to the rubric itself will be flagged and explained in the revision notes, not folded in quietly.

Current scores (v1.0)

Columns are the five dimensions: State (persistence), Depth (adaptation), Feedback (integration), Archetype (differentiation), and Health (relationship tracking).

Tool
State
Depth
Feedback
Archetype
Health
Total
Qualia
8
9
7
9
7
40/50
Gong
7
3
5
2
8
25/50
Crystal Knows
6
4
3
8
2
23/50
Humantic AI
6
4
3
9
1
23/50
Outreach
5
3
5
2
6
21/50
Mem0
8
2
7
1
1
19/50
Contextra
provisional — public docs inaccessible at time of scoring
4
4
3
3
3
17/50
Lavender
4
5
4
2
1
16/50
Clay
5
3
2
2
2
14/50
DIY / custom
3
3
2
2
1
11/50

Where Qualia Falls Short on Its Own Benchmark

We scored ourselves 40 out of 50. Here's exactly where the other ten points are, stated as plainly as we'd state anyone else's gaps.

Multi-language adaptation
partial credit within Adaptation Depth and Archetype Differentiation
The current model is English-optimized. Adaptation quality measurably drops for non-English recipients — not a rounding error, a real degradation. We don't have a shipped fix yet. This is the single largest documented gap in the product.
Cold-contact profiling
partial credit within Recipient State Persistence and Relationship Health Tracking
With little or no prior interaction history, the recipient model has less to work with, by construction — the system gets meaningfully better the more history it accumulates, which means it's meaningfully weaker on a first message to a stranger than on the fiftieth message to a known contact. We consider this an honest structural limitation of relationship-state modeling generally, not unique to us, but we haven't built a mitigation for the cold-start case specifically.

Feedback Integration is real but shallow. The system updates reliably on explicit signal — edits, replies. It's less reliable on implicit signal: silence, delayed response, a tone shift in a reply that doesn't directly reference the previous message. We score ourselves a strong but not perfect mark here for that reason.

What we are not claiming. We are not claiming parity across every relationship type, every language, or every contact-history depth. The 40/50 is the honest number for the product as it exists today, not the product as we intend it to exist.

Roadmap against these gaps: non-English adaptation quality is the priority for the next revision cycle. We're not committing to a ship date publicly until we have real data to back one — we'd rather under-promise here than repeat the kind of confident claim this whole benchmark exists to push back against.

Changelog

v1.0 — June 19, 2026. Initial publication. No prior revisions. Scores are reviewed and revised each quarter; the next update is scheduled for September 17, 2026.

Keep reading
The problem
Recipient Blindness
The solution
Recipient Intelligence
Build with the API
Developer Docs