# The Crossing — Complete Technical Readout
**Instrument v0.5.2 · scoring engine v0.6.1 · readout v1.4 — a first beta-data revision (23 statements) · status: pre-pilot (cognitive-interview stage) · September 2026**

*Revised September 27, 2026: reviewer wording withdrawn; nothing else in this readout changed.*

*This document contains the complete instrument — every question, verbatim — and the full mechanics behind it: what the test claims, how it's built, how it's scored, what evidence supports each design decision, and what remains unproven. It is written to be scrutinized. Section 9 tells you, the reviewer, exactly what feedback we're asking for.*

---

## 1. What this test claims — and refuses to claim

The Crossing is a consumer Enneagram typing instrument: 68 screens, about 9 minutes, producing a **suggested type** with an explicit **confidence band**. Its claims are deliberately modest, because the honest ceiling for this task is lower than the industry pretends:

- It does **not** claim a fixed accuracy percentage. (The best-validated comparable instrument — the Stanford/Essential Enneagram paragraph test, validated on 970 participants against independent expert typing and/or post-education type reassessment — published overall criterion agreement of **κ≈.53**, with per-type examples of 66% (Type 1) and 68% (Type 9). A leading commercial competitor advertises ">95% accuracy"; that figure appears nowhere in its own published statistics. We benchmark against the former and refuse to imitate the latter.)
- Every result ships with a band — **CLEAR / LIKELY / CLOSE / TOSSUP** — computed from the actual score gap, and close calls route to a **self-confirmation step** where the taker chooses between their top candidates. The instrument proposes; the person confirms. The taker's confirmation, not our score, is the final record.
- We commit to publishing our own reliability and per-type accuracy statistics as real data accrues, on a public science page — including the numbers that embarrass us.

One construct-level honesty note up front: the Enneagram's nine-type structure itself has *mixed* support in the factor-analytic literature (a 2021 systematic review of 104 samples found instruments often recover fewer than nine clean factors). We treat the nine types as a useful, coherent framework of long tradition — and measure it as carefully as the state of the art allows — without claiming the underlying taxonomy is settled science.

## 2. The measurement problem this design answers

Typing is harder than trait measurement, for four documented reasons — each of which shaped this instrument:

**2.1 The construct is motivation; self-report delivers behavior and self-image.** The Enneagram defines types by *why* (core fear, core desire), but people cannot reliably introspect on abstract motivation. Items like "I am driven by a need to be good" force interpretive work at the comprehension stage (Tourangeau's four-stage response model) — our own field testing of a 151-item predecessor produced exactly the predicted failures: confusion, clarification requests, long response times. **Design answer:** every item is a concrete, observable experience or behavior chosen as a *downstream marker* of the motivation ("After I make a small mistake, I keep replaying it in my head" is the fingerprint of the Type 1 engine, requiring no introspective theory from the respondent).

**2.2 Acquiescence bias — the documented killer of Type 9 measurement.** Agreeable respondents tend to agree with items regardless of content. Under naive scoring (sum each scale), a yes-saying respondent inflates *every* scale — and whichever scale's items are most broadly endorsable wins. Our predecessor instrument mistyped a known Type 9 as a Type 1 in field testing: the classic signature, since duty/order items are both socially desirable and universally endorsable, and Type 1 has repeatedly shown positive associations with conscientiousness across Enneagram–Big Five studies. **Design answers:** (a) cost-side and attention-side item wording that avoids broadly endorsable virtue statements (§3.1); (b) equal-exposure, desirability-balanced forced-choice blocks, including a dedicated 9/1/5 block, providing head-to-head evidence that blanket agreement cannot inflate; (c) a flat-profile flag that treats undifferentiated response patterns as low signal. (An earlier draft claimed within-person mean-centering as the acquiescence fix; centering cannot change a person's own ranking. The idea is now retired entirely: future population percentiles compare each raw scale against its own reference distribution — within-person pattern and cross-person percentile stay separate concepts, per §5.3.)

**2.3 Social desirability.** Some type content is simply nicer to endorse. **Design answers:** items written at the *cost* end of each pattern (the Type 1 items are about replay-loops and inability to rest, not about "doing things right"); triad options desirability-balanced within each block; an idealized-responding marker (V2) — inert alone as an internal QA flag, and capping the confidence band only in combination with a uniform-high response pattern.

**2.4 Look-alike types.** Adjacent and structurally similar types produce the documented mistype patterns (9↔1, 9↔2, 1↔6, 3↔1, 6↔9, 4↔5…). A Likert scale measures each type's signal but is weak at *discriminating* near neighbors. **Design answer:** the mixed format — Likert core for measurement, forced-choice triads for discrimination, deployed as tie-breakers exactly where the literature says forced-choice earns its keep.

## 3. Design decisions and their evidence basis

**3.1 Item-writing rules (every item, no exceptions).** First person, present tense; one idea per item (no double-barrels); ≤20 words; Flesch-Kincaid grade ≤8 as an authoring target (computed with the Python `textstat` implementation — named because FK values vary by syllable-counting method; Gate 1 comprehension is the binding test); no negations, no idioms, no vague quantifiers; concrete and experiential rather than abstract; sampled across automatic-attention, protective-strategy, and characteristic-cost expressions of each type — an earlier all-cost-side rule was deliberately rebalanced so healthier members of a type still register. Basis: the scale-development canon (Clark & Watson 1995/2019; DeVellis; Krosnick's questionnaire-design research) and the survey-comprehension literature (Tourangeau; Willis's cognitive-interviewing methodology). A well-written Likert item should take a few seconds; our Gate-1 pass bar is median ≤10 seconds and ≥80% comprehension-as-intended in think-aloud interviews.

**3.2 Response format: 5-point, fully labeled agreement scale.** The scale-points literature (Krosnick & Presser; Preston & Colman; Lozano et al.) tends to show improvements up through roughly 5–7 response categories, with diminishing returns beyond; findings on labeling favor full labels, though not uniformly. Five points permits clean full labeling on mobile. A labeled midpoint is included.

**3.3 Scale length: 6 items per type — the reliability math.** Coefficient alpha is a function of item count (k) and average inter-item correlation (r̄): α = k·r̄/[1+(k−1)·r̄]. At a realistic post-refinement r̄ of .30: 3 items → α≈.56 (inadequate for classification); **6 items → α≈.72**; 8 → ≈.77; 10 → ≈.81. The step from 3 to 6 is the steep part of the curve; beyond ~8, gains are marginal while fatigue costs grow. Nine-way classification is *more* error-sensitive than dimensional scoring (the true-highest scale must survive measurement noise in all nine scales; the MBTI analog shows only ~65% same-type retest agreement even with reliable continuous scales), which is why we sit at 6 rather than 3. Total: 54 core items.

**3.4 Test length: 68 screens, ~9 minutes.** Respondent burden and satisficing generally increase as questionnaires lengthen — straight-lining and speeding rise with fatigue, corroding the very inter-item correlations the reliability argument depends on — so we target a sub-10-minute consumer experience. An 80+ item version was evaluated and rejected on this reasoning: past ~8 items/scale, projected reliability gains are marginal while burden costs grow.

**3.5 The triads are tie-breakers, not a standalone measure.** Simulation work on Thurstonian-IRT forced-choice designs (Bürkner, Schulte & Holling 2019) documents serious reliability and ipsativity limitations for small designs with few traits; operational forced-choice instruments use an order of magnitude more blocks. Twelve triads therefore do not attempt to score nine types — but as a **complete balanced pair design** (a Steiner triple system: all 36 type-pairs, each exactly once), they deliver exactly one desirability-balanced head-to-head contest for every possible pair of finalists, which is precisely and only what the tie-breaking role requires.

**3.6 Direct keying, no reverse-worded items — a documented tradeoff.** Classic advice mixes reversed items to disrupt acquiescence; strong recent evidence (Zhang, Noor & Savalei 2016; van Sonderen et al. 2013; Weijters et al.; Schriesheim & Hill 1981) shows reverse-worded items confuse respondents, create method artifacts, and lower reliability. Since comprehension failure was our predecessor's primary field failure, we key all items directly and control acquiescence through cost/attention-side wording, equal-exposure forced-choice blocks, and response-pattern flags instead. This decision is flagged for re-evaluation at pilot if response-pattern flags spike.

**3.7 Identical administration for every taker.** No adaptive routing. This costs some per-screen efficiency (adaptive tests extract more information per item) but buys the thing our validation model requires: fully comparable item-level data across all takers, enabling published item statistics and honest per-type accuracy reporting.

## 4. The complete instrument — every question, verbatim

*Extracted directly from the canonical item bank (v0.5.2). Administration: one item per screen, auto-advance; **response labels, exact and always visible: Strongly disagree · Disagree · Neither agree nor disagree · Agree · Strongly agree**; one step of back-navigation permitted (changes logged); **both the order of the twelve triads and the statement positions within each triad randomized at runtime with independently logged seeds** — so no type is confounded with fatigue position; fixed interleave for Part A (three counterbalanced development forms through Gate 2, one locked production form after); autosave; response time logged per item. Two legs: Part A (54 core + 2 validity), then Part B (12 triads).*

### 4.1 Part A — The 54 core items (verbatim, bank v0.5.2; six per type)

**Type 1 — The Compass** *(inner critic · error radar · standards enforcement · compressed resentment · self-audit voice · earned rest)*
- A1.1 — After I make a small mistake, I keep replaying it in my head.
- A1.2 — Small errors nag at me until they're fixed — even ones nobody else would notice.
- A1.3 — I redo other people's work so it gets done the right way.
- A1.4 — My frustration with sloppy work leaks into my tone before I can stop it.
- A1.5 — A voice in my head points out what I could have done better.
- A1.6 — Rest feels like something I have to earn by finishing everything first.

**Type 2 — The Tender** *(needs radar · over-giving · asking difficulty · recognition sting · usefulness strategy · refusal difficulty)* — four items revised in v0.5.2 (cross-loading)
- A2.1 — I notice what people need before they say anything.
- A2.2 — I offer help before anyone asks for it.
- A2.3 — I find it hard to ask for help, even from people close to me.
- A2.4 — I quietly keep track of all I do for people, and it stings when nobody notices.
- A2.5 — Being needed is how I know I belong.
- A2.6 — I say yes because I cannot stand for someone to feel let down by me.

**Type 3 — The Navigator** *(goal conversion · image adaptation · restlessness · pace impatience · worth-through-results · motion as comfort)* — unchanged
- A3.1 — I turn most things into a goal I can win or finish.
- A3.2 — I adjust how I come across depending on who is watching.
- A3.3 — After I reach a goal, I move to the next one instead of celebrating.
- A3.4 — I get impatient when people or meetings slow me down.
- A3.5 — I feel like my value depends on what I accomplish.
- A3.6 — I stay busy because stopping makes me uneasy.

**Type 4 — The Wayfinder** *(differentness · feeling-mining · comparison-to-missing · longing pull · identity stakes · envy sting)*
- A4.1 — I feel different from other people, even in groups I belong to.
- A4.2 — When a strong feeling arrives, I turn it over for hours, searching for what it means.
- A4.3 — I often feel a pull toward a life that would feel more true than this one.
- A4.4 — I long for what is absent more than I enjoy what is present.
- A4.5 — I feel my differences from people more than my similarities.
- A4.6 — Other people's happiness can sting when mine feels far away.

**Type 5 — The Cartographer** *(privacy · resource accounting · prepare-before-engage · observer stance · needs minimization · competence-before-audience)*
- A5.1 — I keep most of my thoughts and feelings to myself.
- A5.2 — Before I say yes to plans, I work out what they will cost me in energy.
- A5.3 — I hang back and study a new situation from the edge before stepping in.
- A5.4 — In emotional moments, I step back and observe instead of reacting.
- A5.5 — I keep my needs small so I do not depend on others.
- A5.6 — I practice things in private until I am sure of them — being watched while I learn drains me.

**Type 6 — The Watchkeeper** *(risk simulation · verification · backup-seeking · trust testing · durable loyalty · decision doubt)*
- A6.1 — When plans are made, I picture what could go wrong.
- A6.2 — I double-check information before I rely on it.
- A6.3 — I feel better about a decision when someone I trust agrees with it.
- A6.4 — I test whether people will still be there when things get hard.
- A6.5 — I stay loyal past the point that serves me, because leaving feels riskier than staying.
- A6.6 — I second-guess decisions after I have made them.

**Type 7 — The Voyager** *(anticipation · heaviness exit · option-foreclosure · choice-as-loss · pain reframing · anticipation-beats-experience)* — full six-item rebuild in v0.5.2 (weak coherence)
- A7.1 — My mind is already tasting the next thing while I'm still in the middle of this one.
- A7.2 — I change the subject when talk turns painful — even when I know I shouldn't.
- A7.3 — I feel trapped the moment my options start to narrow.
- A7.4 — Choosing one thing feels like losing all the others.
- A7.5 — I find the upside so fast that the painful part never gets a turn.
- A7.6 — The best part of anything is usually right before it starts.

**Type 8 — The Breakwater** *(directness · counter-control · takeover reflex · impact blindness · vulnerability armor · power radar)*
- A8.1 — I say what I think, even when it will start a conflict.
- A8.2 — When someone tells me what to do, I push back.
- A8.3 — I step in and take charge when no one else does.
- A8.4 — People tell me I come on stronger than I realize.
- A8.5 — Showing weakness feels dangerous to me.
- A8.6 — Walking into a room, I quickly sense who really holds the power.

**Type 9 — The Harbor** *(merging · slow no · want suppression · conflict cost math · self-forgetting · routine inertia)*
- A9.1 — I go along with other people's plans instead of saying what I want.
- A9.2 — My attention drifts to comfortable distractions when something needs deciding.
- A9.3 — When people ask what I want, I say I am fine with anything.
- A9.4 — Raising a problem feels harder to me than living with it.
- A9.5 — I lose track of what I want when others have strong opinions.
- A9.6 — I keep doing things the familiar way even when a better way is offered.

**Validity items (2, interleaved):** V1 attention check ("Choose 'Agree' for this one."). V2 idealized-responding marker ("I have never said anything unkind about anyone.") — **v0.4 role change: V2 alone is a soft flag only (logged, disclosed in the report's typed-you appendix); it caps the band only in combination with a uniform-high response pattern.**

### 4.2 Part B — The 12 tie-breaker triads (verbatim, bank v0.5.2; every type-pair exactly once)

*The design is a Steiner triple system S(2,3,9) — the affine plane AG(2,3): 12 triads over 9 types in which **all 36 possible type-pairs occur exactly once** and **every type appears exactly 4 times**. This eliminates the missing-pair and duplicate-pair problems structurally: whatever pair of finalists Part A produces, exactly one direct head-to-head contest exists. Three v0.4 blocks were already Steiner blocks and survive verbatim (T1, T2, T3 below). All statements are *designed* to be dignity-balanced — each carrying a virtue and a cost shading, parallel construction wherever possible; that balance is an assumption tested at Gates 1–2, not asserted (Gate 1 direct probe: "which of these three answers makes someone sound best?").*

- **T1 (9 / 1 / 5)** *(carried from v0.4 — the dedicated 9↔1 block; Nine statement revised in v0.5.2)*: (9) I let things settle on their own timing — rushing a decision usually makes it worse. · (1) I fix what is wrong as soon as I see it, even when no one asked. · (5) I gather more information before I let myself be pulled in.
- **T2 (9 / 2 / 4)** *(carried from v0.4)*: (2) I set my own needs aside to look after someone else's. · (9) I keep the group steady by staying easy to be around. · (4) I stay true to my feelings even when they complicate things.
- **T3 (4 / 5 / 6)** *(carried from v0.4)*: (4) I trust my feelings more than any outside opinion. · (5) I trust my own analysis more than other people's advice. · (6) I trust advice from people who have earned it more than my own hunches.
- **T4 (1 / 2 / 3)** — how I earn my place: (1) I earn my place by getting things right. · (2) I earn my place by being there for people. · (3) I earn my place by delivering results.
- **T5 (7 / 8 / 9)** — under pressure *(the Storm-and-Harbor head-to-head)*: (7) When life gets hard, I look for the open door. · (8) When life gets hard, I push straight through. · (9) When life gets hard, I wait for the water to settle.
- **T6 (1 / 4 / 7)** — what jumps out first: (1) What is wrong with a thing jumps out at me first. · (4) What is missing from a thing pulls at me first. · (7) What could happen next excites me first.
- **T7 (2 / 5 / 8)** — how I spend myself: (2) I give my time freely, even when I am running low. · (5) I guard my time closely, even with people I like. · (8) I put my time where I control the outcome.
- **T8 (3 / 6 / 9)** — when I feel best *(the 6↔9 head-to-head)*: (3) I feel best when I am making visible progress. · (6) I feel best when I know the ground is solid. · (9) I feel best when nothing is pulling at me.
- **T9 (2 / 6 / 7)** — when someone struggles: (2) I move in and help carry it. · (6) I help them plan for what comes next. · (7) I help them find the bright side.
- **T10 (3 / 4 / 8)** — how I want to be seen *(Three statement revised in v0.5.2)*: (3) I want to be seen as impressive — being merely useful stings a little. · (4) I want to be seen as one of a kind. · (8) I want to be seen as strong.
- **T11 (3 / 5 / 7)** — what gives me energy: (3) Finishing things gives me energy. · (5) Understanding things gives me energy. · (7) Starting things gives me energy.
- **T12 (1 / 6 / 8)** — where authority sits *(One statement revised in v0.5.2)*: (1) I hold myself to the rules even when no one is watching — and judge myself hard when I slip. · (6) I want to know who is in charge and whether they have earned it. · (8) I answer to myself first, whoever is in charge.

*(The v0.4 blocks not in the Steiner system — including the old 2/3/7 — are retired; their discriminations are inherited by T4, T9, and T11.)*

*Triad watch list (Gate 1 priority, desk edits suspended): T7's Eight statement, T10's Four statement, and T12's One/Eight pairing carry possible desirability asymmetry — probe directly; a >60% MOST-rate at Gate 2 is a review flag for desirability, interpretation, or base-rate effects, never an automatic rewrite.*

## 5. The scoring pipeline (engine v0.6.1) — integer logic, with test cases

**5.1 The steps.** Server-side only; coded on integer six-item sums (6–30); means are display-only; all thresholds provisional pending Gate 3b calibration.
1. **Validity flags (crisp definitions):** V1 answered anything but "Agree" → cap CLOSE. "Strongly endorsed" V2 = Agree or Strongly agree; "uniform-high" = grand mean ≥ 4.25 across the 54 core items; both together = IDEALIZED → cap LIKELY; V2 alone = internal QA flag — logged, never consumer-disclosed, no cap; the taker is told about whatever materially affected their result, never accused over inert flags.
2. **Sums and gap:** per-type sum of 6 items; Δ = top1 − top2 in whole response points.
3. **FLAT flag:** range of nine sums ≤ 4, or ≥5 sums within 2 points of top1 → cap CLOSE, forks always shown. Flatness lowers confidence; it can never raise it.
4. **Decision rule with the runner-up set:** R = every type tied at the second-highest sum (top2 need not be unique); *agreement* = top1 defeats **every member of R** in their shared triads, so the band can never depend on arbitrary internal tie ordering. Δ ≥ 3 → **CLEAR** with full agreement and no flags, else **LIKELY**. Δ = 2 → **LIKELY** with full agreement and no flags, else **CLOSE**. Δ ≤ 1 → tie-breaker mode.
5. **Tie-breaking is a pairwise round-robin.** Candidates = all types within 1 point of top1. Each triad's MOST/LEAST yields a full ordering, so every candidate pair's unique shared triad produces a decisive head-to-head (within-pair ties impossible). Two candidates → single contest, band CLOSE. Three or more → round-robin: a Condorcet winner takes it at CLOSE; a cycle or tied win-count falls to the higher Part A sum if unique (band **CLOSE** — a unique measurement leader existed; the discriminator merely failed to resolve the contest) — otherwise the result is **TOSSUP** and the confirmation forks decide, honestly.
6. **Bands:** CLEAR / LIKELY / CLOSE / TOSSUP ("CLEAR" describes the observed separation; epistemic confidence language is reserved until Gate 3b calibration). Flags only lower bands. The honesty line renders with every result: *"Your band reflects how clearly one pattern separated from the others in your answers — not a probability that the result is correct."*
7. **Confirmation** logged as face-validity/product data, never as an accuracy criterion. **Chart:** raw sums scaled for display, fixed 1→9 order, untouched by tie-breakers. **Versioning:** byte-reproducible including both Part B randomization seeds (block order and statement position).

**5.2 Worked examples — constructed cases, preserved as permanent regression tests.**
*The flat-profile trap:* sums 25, 24, 24, 24, 24, 24, 24, 24, 24. Δ=1 → tie-breaker; range=1 → FLAT; band capped at CLOSE with forks. (The retired z-logic standardized this near-flat profile into false-CLEAR territory; v0.6 treats "everything looks the same" as exactly what it is.)
*The differentiated profile:* sums 27, 25, 23, 21, 18, 15, 12, 9, 6. Δ=2 → the LIKELY path, held there by head-to-head agreement. Larger real separation now always means more confidence, never less.
*The acquiescent Type 9:* blanket agreement inflates every sum, but cost/attention-side items deny Type 1 its virtue subsidy, and triad T1 — where the Nine statement is as dignified as the One and Five statements — decides any residual near-tie on evidence agreement cannot fake.
*The 6-versus-9 near-tie (previously unresolvable):* sums tied at 26 → candidate pair {6, 9} → triad T8 ("solid ground" vs. "nothing pulling at me") delivers the direct contest the old nine-block design structurally lacked.
*The tied-runner-up case:* sums 27 / 24 / 24. Under the prior wording, the band depended on which tied type the software happened to call "top2" — nondeterminism in a system promising byte-reproducibility. Under the R-set rule, Type 5 must defeat *both* tied runners for full agreement; a split downgrades. Identical answers now always produce identical bands.

**5.3 What the taker sees.** The nine-bar chart, the band with its honesty line, the confirmation step, and — in the paid report — the "How we typed you" appendix showing their own top-two, gap, band logic, flags, and runner-up check. Future percentile norming compares each raw scale against its own reference distribution; within-person pattern and population percentile remain two distinct, separately-labeled concepts.

## 6. Validation program — how "reliable" gets earned (v3)

- **Gate 1 — cognitive interviews (current stage).** Willis protocol, 3 iterative rounds × 5; priority on the formal watch-list items (A4.5, A5.2, A6.5, A9.6, the Type 1 work-saturation trio, the Type 9 preference-redundancy trio, and six older flagged items), **all twelve triads** (T6/T7/T10/T12 especially), deliberately recruited counterphobic-leaning Sixes, and a direct desirability probe on every triad ("which of these three answers makes someone sound best?" — a consistent winner is a real desirability problem independent of personality). Pass: ≥80% comprehension-as-intended, median ≤10s/item; twice-failed items rewritten from scratch.
- **Gate 2 — beta pilot (~300; beta evidence only — no accuracy claims).** Item-total r < .30 is a **review flag, not an automatic removal** — a distinct-facet item may run lower, and alpha-maximizing can narrow the construct. Also: α/ω/AIC as measured; completion ≥85%, median ≤10 min; triad MOST-rate balance (>60% → review flag for desirability, interpretation, or base-rate effects — genuine prevalence is not a defect); FLAT prevalence; scale-location comparison across the nine types (an endorsement-difficulty check — systematic differences get flagged for criterion-informed calibration, never desk-corrected); structural analysis **exploratory only**.
- **Gate 3a — product metrics (ongoing, labeled as such):** confirmation agreement = face validity; retest = stability — reported as **both** continuous scale reliability and categorical same-type agreement, because scores can be stable while a hairline top-two switch changes the label.
- **Gate 3b — calibration on independent criterion data.** Blinded structured typing by two raters, **inter-rater agreement reported before adjudication**, plus a type-after-education-before-results arm; recruitment is **criterion-type aware** — a pre-specified minimum of independently criterion-classified cases per type (set by a precision calculation at protocol design; overall cohort ~900+). This gate **tunes** thresholds, flags, and rules, with **confirmatory** structure testing here — never in the sample that explored it.
- **Gate 3c — locked holdout validation.** Bank, engine, thresholds, flags, and tie rules freeze; evaluation runs on criterion-classified cases untouched by any tuning decision (pre-registered train/holdout split), reporting pre-registered metrics: overall accuracy, per-type sensitivity/recall, per-type precision/PPV, specificity where useful, κ, and macro-averages — never one ambiguous "accuracy" number. **Only Gate 3c licenses public accuracy claims.**
- **Long-term docket:** differential item functioning across demographic groups and prior Enneagram familiarity.
- **Factor-structure policy:** fewer than nine clean factors → construct work (cross-loading analysis, content refinement, replication), never heavier classification pressure. Findings publish either way. Norm updates never retroactively alter delivered reports.

## 7. Honest limitations — current, stated plainly

1. **The Crossing is scientifically constructed; its reliability and validity have not yet been established empirically.** The architecture is grounded in established test-construction principles; operational thresholds that lack empirical calibration are explicitly provisional, and no reliability or accuracy statistic for *this* instrument exists yet. That is what the gates are for, and we say so in the product.
2. **Three-to-nine-factor risk.** Factor analysis of Enneagram instruments often recovers fewer than nine clean factors; if our pilot does too, we will report it and respond with construct work — cross-loading analysis, content refinement, replication — never with heavier classification pressure.
3. **Self-report ceiling.** No self-report instrument can fully see around self-image; our confirmation step and per-type accuracy reporting are mitigations, not cures.
4. **Instincts unmeasured.** The tradition's instinctual-subtype layer is not scored; the report teaches it as a self-locate section, labeled as unscored, until an honest measure exists.
5. **The tie-breaker layer is deliberately thin.** Twelve triads discriminate; they do not measure. Their contribution is bounded by design (band effects and near-tie resolution only).

## 8. Benchmarks — the honest comparison table

| Instrument | Length | Published performance |
|---|---|---|
| Essential/Stanford Enneagram (paragraph method) | 9 paragraphs + discriminators | overall criterion agreement κ≈.53 vs. independent expert typing and/or post-education reassessment (n=970); ~4-week retest κ≈.59; published per-type examples: Type 1 66%, Type 9 68% |
| RHETI 2.5 | 144 forced-choice pairs, ~40 min | peer-reviewed evidence (Newgent et al. 2004): adequate internal consistency, mixed construct validity; factor analyses recover fewer than nine factors; frequently-cited categorical retest figures are undergoing our source audit before any public-facing use |
| iEQ9 | ~175 adaptive items | vendor-published per-type α .73–.84; advertised ">95% accuracy" absent from its own statistical documentation |
| MBTI (classification analog) | 93 items | ~65% same four-letter type on retest — the cautionary tale for categorical scoring |
| **The Crossing v0.5.2** | 68 screens, ~9 min | projected per-scale α≈.72–.76 at r̄=.30–.35 (projection, not measurement); classification accuracy TBD at Gate 3b, to be published per type with a confusion matrix |

## 9. For the reviewer — what we're asking you to attack

1. **Read the 54 items cold.** Flag any item that is ambiguous, double-barreled, above an 8th-grade read, or answerable two different ways by the same person. Tell us which word stopped you.
2. **Time yourself.** Any item that takes you >10 seconds is a finding.
3. **Attack the triads.** In each block, is one option clearly the "good" answer? If you can rank the three by social desirability easily, the block needs rebalancing — tell us the ranking you see.
4. **Attack the scoring.** Do the integer worked examples (§5.2) convince you? Can you construct a response pattern that fools the decision rule (§5.1) — or a triad ordering that produces an unhandled edge case?
5. **Attack the claims.** Is any sentence in this document a claim we haven't earned? That's the one we most want to hear about.

Every flag you raise feeds Gate 1 directly. This is the review process working as designed.

*Readout v1.1 change note: this document was corrected in four places — the acquiescence mechanism (§2.2, §5), the confidence algorithm (§5, now engine v0.5 with raw gaps, flat-profile detection, and candidate-specific tie-breaks), the validation criterion structure (§6, self-confirmation reclassified; independent Gate 3b added), and benchmark language (§1, §8, κ-based). The constructed failure cases are preserved in §5.2 as permanent regression tests.*

*Readout v1.2 change note: Part B rebuilt as a complete balanced pair design (12 triads, all 36 pairs exactly once); engine v0.6 recoded on integer sums with pairwise round-robin tie-breaking; the top band renamed CLEAR; edge cases, response labels, back-navigation, and position-randomization fully specified; validation additions (pre-adjudication rater agreement, r<.30 as review flag, exploratory/confirmatory split, continuous+categorical retest, DIF docket); and remaining internal contradictions and overstated claims corrected throughout.*

*Readout v1.4 change note (September 2026 — first beta-data revision): the first beta cohort's item-level data drove a 23-statement revision — instrument v0.5.1 → v0.5.2. Twenty Part-A items (Type 7 rebuilt in full; Type 2 ×4; Types 1, 4, 5, 6, 9) and three tie-breaker statements (T1 Nine, T10 Three, T12 One) were rewritten to correct an endorsement-difficulty imbalance across scales, weak Type-7 coherence, items cross-loading on look-alike types, and three tie-breakers winning on social desirability. Statement text only — item IDs, types, keying, facet structure, the 54-item count, the Steiner triad design, and scoring engine v0.6.1 are unchanged; confidence bands are unchanged (the frozen longitudinal metric). The version-segmented dashboard renders the verdict at ~25 v0.5.2 sessions, published either way. Full old→new change log with the indicting statistic per item: the canonical item bank (crossing-v2-item-bank-v0-5.md).*

*Readout v1.3 change note: the tied-runner-up set R closes the band-nondeterminism edge case; the round-robin Part-A fallback band is fixed at CLOSE; T1–T12 block order is runtime-randomized alongside statement positions; V2-alone becomes internal-only; "dignity-balanced" is restated as a tested design intent with a direct desirability probe at Gate 1; Gate 3 splits into 3b calibration and 3c locked holdout with pre-registered metrics and criterion-type-aware recruitment; the readability target names its implementation; and the version-line and criterion-language contradictions — casualties of a crashed patch run — are corrected under a new verification-sweep discipline.*

---
*Methodological sources referenced: Clark & Watson (1995, 2019); DeVellis, Scale Development; Krosnick & Presser (2010); Tourangeau, Rips & Rasinski (2000); Willis, Cognitive Interviewing (2005); Livingston & Lewis (1995); Soto & John (2017, BFI-2); Bürkner, Schulte & Holling (2019); Zhang, Noor & Savalei (2016); Daniels & Price (Essential Enneagram validation, n=970); Newgent et al. (2004, RHETI); Hook et al. (2021, J. Clinical Psychology systematic review, 104 samples).*
