The thesis
We don't ask what you do. We ask what you're protecting.
Most personality tests ask about behavior — are you outgoing, are you organized. The trouble is that the same behavior can come from nine different places. Two people can both avoid an argument: one to keep the peace that holds their world together, one because the argument might expose a weakness they guard.
The Enneagram’s real subject is the why underneath the what — but you can’t just ask people their deepest motivation, because almost nobody can answer that honestly on a Tuesday. So The Crossing works in the layer between: behavioral moments that contain the protective logic.
Three of our actual questions:
“Being needed is how I know I belong.” — not “are you helpful,” but the heart-level engine: being needed as proof of belonging.
“When people start expecting things from me, my first instinct is to pull back.” — not “are you independent,” but the reflex underneath: expectation felt as a claim, and a step back to guard what’s yours.
“Raising a problem feels harder to me than living with it.” — not “do you avoid conflict,” but the exact cost-math that makes a person tolerate what hurts them.
Fifty-four questions like these — six for each of the nine patterns, mixing what your attention does automatically, the strategy you reach for, and the cost you pay when it runs too long. Then twelve quick either-or rounds built for one purpose: telling apart the patterns that look alike. (Every one of the 36 possible pattern-pairs gets exactly one direct head-to-head — a mathematically complete design. We can prove that sentence; the full readout does.)
About 10–12 minutes. No trick questions. Nothing you need a theory to answer.
Construction
Built like an instrument, not a quiz.
- The questions follow rules. First person, present tense, one idea each, under twenty words, readable at an eighth-grade level, no double meanings. Every item survived (or will face) live think-aloud testing with real people before the test is finalized.
- The wording fights the two classic failure modes. Personality tests fail when nice-sounding answers score points (social desirability) and when agreeable people agree with everything (acquiescence). Our items are written so that no pattern gets a virtue subsidy — and the either-or rounds are built so that no option is “the good one.”
- Everything is logged, nothing is adaptive. Every taker gets the same instrument, so every taker’s data can honestly be compared — the foundation for the statistics we intend to publish.
Scoring
The part most tests hide is the part we show.
Your fifty-four answers become nine scores — one per pattern. If one pattern clearly stands apart, that’s your suggested type. If two or three run close, the either-or rounds act as tie-breakers: your finalists meet in their direct head-to-head contests, and the winner takes it. And if your answers genuinely don’t separate — it happens — we say so instead of faking a verdict, and you make the final call between portraits.
Every result ships with one of four bands:
| Band | What it means |
|---|---|
| Clear lead | Your top type scored clearly ahead of the others, and your 'most like you / least like you' answers agreed. |
| Moderate lead | Your top type came out ahead, but not by much, or some of your 'most like you / least like you' answers pointed elsewhere. |
| Close call | Your top types scored close together, so read your runner-up too. |
| Split | Your answers came out even between two or more types. |
And beneath every band, this sentence, always:
“Your band reflects how clearly one pattern separated from the others in your answers — not a probability that the result is correct.”
Your paid report includes a “How We Typed You” appendix: your own top two patterns, the gap between them, which rule fired, and the honest check to run if the result doesn’t sit right. If you conclude we called it wrong, your confirmation — not our score — is the record.
Version history
The change log — every version, in plain words.
We said the change log stays public, including the parts that sting. Here it is: what each revision found, what we changed, and how we’ll know if it worked. No accuracy claims — these are construction improvements, judged on our own version-labeled data.
Instrument v0.5.10 · engine v0.6.4September 30, 2026
Question update (instrument v0.5.10): we rewrote one tie-breaker choice so it aims more directly at its type's lasting pattern.
Instrument v0.5.9 · engine v0.6.4September 30, 2026
Question update (instrument v0.5.9): we rewrote 12 statements and 4 tie-breaker choices that were picking up other types, a hard stretch of life, or wording people tended to avoid. Each now aims more directly at its type's lasting pattern.
Instrument v0.5.8 · engine v0.6.4September 30, 2026
Engine v0.6.4: each result label now comes with a plain explanation of what it means, and if your 'most like you / least like you' answers leaned toward a type other than your top two, your result shows it.
Instrument v0.5.7 · engine v0.6.3September 27, 2026
Engine v0.6.3: when your answers don't separate the leading types, we now show them together and help you tell them apart, instead of picking the one with the lowest number. We also removed wording that claimed more certainty than we have.
Instrument v0.5.7 · engine v0.6.2September 2026
We restyled the tie-breaker questions. No statement changed.
What we found
- Each tie-breaker screen had its scene-setting phrase in small caps above the question, and on five of the twelve screens the three answer choices repeated the same opening words three times in a row.
What we changed
The phrase now reads as one plain sentence beneath the question, and those five screens show its shared opening words once instead of three times. What a person reads is exactly the same words as before — nothing about the test's content or scoring changed.
How we’ll judge it
This is a presentation change to the tie-breaker screens themselves, so it starts its own 25-clean-session window rather than continuing the previous version's.
Instrument v0.5.6 · engine v0.6.2September 2026
We removed one of the three reading checks. No statement changed.
What we found
- Three reading checks in a row of 58 screens was more than the test needed: the oldest one ("Choose 'Agree' for this one") had already stopped affecting anyone's result, and takers told us the checks felt repetitive.
What we changed
That one screen is gone: 69 screens instead of 70. The two remaining reading checks, every statement, and the scoring are exactly as they were in the previous version, so results from both versions are read together.
How we’ll judge it
Nothing new to judge — this version continues the previous one's pre-registered predictions on the same 25 clean sessions.
Instrument v0.5.5 · engine v0.6.2September 2026
Our third beta revision: fourteen statement changes aimed at the types the test was mixing up.
What we found
- Five was being suggested more often than people agreed with: several of its statements described quiet, modest habits that many types share, instead of what being a Five actually costs.
- Seven was almost never suggested: its statements described the pattern as a flaw, and Sevens don't see it that way, so they didn't agree with them.
- Some One and Three statements measured the same thing — finishing and staying busy — so the test often couldn't separate those two types, and one tie-breaker line pulled the wrong way.
- Several Six statements described everyday caution almost anyone would agree with.
What we changed
We rewrote twelve core statements and two tie-breaker statements so each one names something that pattern pays and a look-alike type wouldn't. The Six statements now also include the Six who charges straight at what scares them, not only the cautious Six. Nothing else changed: the same number of questions, the same attention checks, and the same scoring math as the previous version.
How we’ll judge it
We pre-registered the predictions before building: Five should be suggested less often and agreed with more, Seven should become reachable, One and Three should separate, and the Six statements should hang together better — without weakening the types that already worked. We judge on the first 25 clean sessions of this version and publish what we find either way.
Instrument v0.5.4 · engine v0.6.2September 2026
Our second beta revision: thirteen statement changes, and a fairer attention check.
What we found
- Several statements were still easier for a look-alike type to agree with than for their own type — quiet-virtue wording a private person of almost any type could endorse.
- The Nine's tie-breaker statements kept losing on tone rather than on fit — written as waiting and stillness, they read as passive next to the other types' lines.
- Our single attention-check question was catching the wrong people: the takers it flagged answered everything else as carefully as anyone — it was reading personality, not attention.
What we changed
We rewrote six core statements and five tie-breaker statements so each pattern confesses its own cost — not a cost anyone could claim. We replaced the single attention check with a pair of plainly-labeled reading checks; the old check stays on screen and logged so the two eras stay comparable, but it no longer affects anyone's confidence band. Only missing both reading checks now does — and missing one does nothing. The scoring engine's type-calling math is untouched; the engine number moves (v0.6.2) because the flag rule changed.
How we’ll judge it
We pre-registered the predictions before shipping: the look-alike leaning should ease, the Nine's own statements should start winning their tie-breakers, and the new flag should mark genuinely careless sessions — lower internal consistency than clean ones — or it gets demoted to a log-only note. We adjudicate on the first 25 clean sessions and publish what we find either way.
Instrument v0.5.3September 2026
We added a one-line context cue above each tie-breaker — no wording changed.
What we found
- Some tie-breakers were written against a situation the screen never actually showed — you met “I help them plan for what comes next” with no “them” in front of you — so the choice arrived without its frame.
What we changed
Each tie-breaker now shows a short context line above the choice (for example, “SOMEONE CLOSE TO YOU IS STRUGGLING”). The statements themselves are byte-identical to v0.5.2 — this is a display change. But because it changes what you see at a scoring moment, we version-stamp it and track v0.5.3 separately on the dashboard.
How we’ll judge it
We'll watch whether the choices that depended on that missing context shift now that it's on screen — and publish what we find.
Instrument v0.5.2September 2026
Our first beta cohort's data drove a 23-statement revision.
What we found
- An endorsement-difficulty imbalance across the scales — some statements were too easy to agree with, others too hard — which bent the raw scores.
- Weak coherence on the Type 7 scale: its statements weren't hanging together as one measure.
- Several statements that leaned toward a look-alike type about as much as their own.
- Three tie-breaker statements that tended to win on sounding good rather than on fit.
What we changed
We rewrote 20 core statements and 3 tie-breaker statements to price each pattern more honestly and to pull look-alike types apart. The test's structure, its scoring engine, and its confidence bands are unchanged — we sharpened the wording, not the cutoffs.
How we’ll judge it
We'll watch the next roughly 25 people on a version-labeled dashboard — whether the imbalance flattens, whether any tie-breaker dominates, and whether fewer people have to correct their type — and publish what we find either way.
Validation
Where the science stands — right now.
A test earns trust in stages. Here is exactly where The Crossing is in that process, updated as each gate completes. We will not claim what a gate hasn’t unlocked.
Gate 1 — Cognitive interviewsIN PROGRESS
Real people take the test out loud while we watch every question earn its place. Unlocks: Final wording lock.
Three iterative rounds of five people each take the test out loud — a “think-aloud” interview — so we can see exactly where a question is misread, slow, or answerable two different ways. Every item has to be understood as intended by at least 80% of people, in under about ten seconds, before the wording is locked.
Gate 2 — Beta pilot (~300 takers)NOT STARTED — you can be one of the 300
First real measurement: reliability, timing, item performance — labeled beta evidence. Unlocks: Published beta reliability numbers.
A pilot of roughly 300 takers gives us our first real measurements — how internally consistent each of the nine scales is, how long the test takes, how often people finish, and how each item and tie-breaker actually performs. All of it published labeled as beta evidence, never as accuracy.
Gate 3a — Stability trackingBEGINS AT LAUNCH
Do results hold on retake? Does the suggested type feel right?. Unlocks: Published stability + face-validity figures.
Once the test is live, we track whether people get the same result on retake (stability) and whether the suggested type feels right to them (face validity) — reported as both a continuous number and same-type agreement, because a score can be stable while a hairline top-two switch changes the label.
Gate 3b — Independent criterion studyNOT STARTED
Blinded expert typing, compared against our results, used to calibrate. Unlocks: Calibrated confidence bands.
Independently trained raters type a set of people without seeing our result; we compare, and use the comparison to calibrate our confidence bands against real criterion data — always on a separate sample from the one we later validate on.
Gate 3c — Locked holdout validationNOT STARTED
The frozen, finished test evaluated on people untouched by any tuning. Unlocks: The only gate that licenses accuracy claims — per-type numbers, full confusion matrix.
The frozen, finished test is evaluated one final time on people untouched by any tuning, with the metrics pre-registered in advance: per-type sensitivity and precision, the full confusion matrix, and κ. This is the only gate that lets us publish accuracy numbers.
The numbers
The table we intend to fill in.
Here is every number we plan to publish, and its honest current value:
| Measure | Current value | Arrives at |
|---|---|---|
| Per-pattern reliability (α/ω) | Not yet measured | Gate 2 |
| Completion rate & median duration | Not yet measured | Gate 2 |
| Retest stability (continuous & same-type) | Not yet measured | Gate 3a |
| Per-type accuracy — sensitivity and precision | Not yet measured | Gate 3c |
| Full 9×9 confusion matrix, κ | Not yet measured | Gate 3c |
This is the full set of measurements a serious instrument stands behind. As each study completes, its results publish here — complete, and per type.
How we compare
The honest benchmarks.
The best-validated comparable instrument — an academically published paragraph-based Enneagram test studied on 970 participants against independent expert typing and post-education reassessment — reported overall agreement of κ≈.53 (a statistic that corrects for lucky guesses), with published per-type examples of 66% and 68%. A separate widely used 144-question instrument shows adequate internal consistency and mixed structural evidence in peer review. The famous four-letter test, the cautionary tale of categorical typing, shows only about 65% of people receiving the same type on retake. And one widely marketed Enneagram test advertises “95%+ accuracy” — a figure that does not appear in its own published statistical documentation.
Those are the real numbers in this category. They are the real bar in this category — and the one The Crossing is built to clear properly: our accuracy figures arrive from Gate 3c, measured against a locked instrument and published per type.
Citations & sources
- Daniels, D. & Price, V. — the Essential Enneagram validation study (n=970): overall criterion agreement κ≈.53 against independent expert typing and/or post-education reassessment; ~4-week retest κ≈.59; published per-type examples Type 1 66%, Type 9 68%.
- Newgent, R. A., et al. (2004) — peer-reviewed evaluation of a 144-item forced-choice Enneagram instrument: adequate internal consistency, mixed construct validity; factor analyses recover fewer than nine factors.
- The four-letter type indicator (categorical-scoring analog): approximately 65% of people receive the same four-letter type on retest — the cautionary tale for categorical scoring.
- Hook, J. N., et al. (2021), Journal of Clinical Psychology — systematic review of 104 Enneagram samples: reliability results vary and analyses do not consistently recover nine clean factors.
- The “95%+ accuracy” figure is drawn from a widely marketed Enneagram product’s advertising and is kept anonymous here; it does not appear in that product’s own published statistics. (Named-comparison policy pending attorney review.)
Our promises
What we will never do — in writing.
- Accuracy numbers come only from our own validation studies (Gate 3c). Then: full publication, per type.
- Every result carries its confidence band and the honesty line. No exceptions, no premium tier that buys false certainty.
- Your confirmation is the record. If you conclude we typed you wrong, the report belongs to the type you claim.
- No celebrity typings. Ever. Typing people we’ve never assessed is speculation dressed as authority — the exact move this brand exists to refuse.
- No “rarest type” statistics until our own norms exist; then, published as ours, with the sample described.
- Delivered reports are never retroactively altered by later norm or scoring updates.
- The full methodology stays public — every question, every rule, downloadable below.
- The change log stays public — every version, every correction, including the embarrassing ones.
- Report access is permanent — never gated behind any membership.
- When we don’t know, we say so. “Not yet measured” is an answer we will always be willing to publish.
Objections
Questions the skeptic asks.
Is the Enneagram itself scientific?
Why should I trust a test that admits uncertainty?
What if the test gets me wrong?
Can I read the actual methodology?
The complete instrument and methodology, unabridged. Every published version stays online at its own address.
10–12 minutes, and it tells you when it’s a close call.