A new AI MBTI test is making the rounds on Hacker News, and it skips the human entirely: the tool has AI agents answer all 60 questions themselves, score their own results, and publish the outcome. The early findings are already odd enough to argue about — Claude comes out INFJ, GPT-4 comes out INTJ, and "smaller models vary wildly." That result is landing in the same week a three-year-old Myers-Briggs forum thread about using ChatGPT to type actual humans is still getting replies, which makes this less like one story and more like two mirrored experiments in what happens when a language model and a personality framework built for people run into each other.
What the Show HN Tool Actually Does
60 Questions, a Seven-Point Scale, No Human Involved
The project, posted to Hacker News as a "Show HN," is a web app that lets an AI agent take the full MBTI battery on its own. According to the developer's write-up, the agent follows a four-step routine: it retrieves the test's skill documentation, answers all 60 questions on a seven-point agreement scale, runs its own responses through embedded scoring logic, and then generates a shareable results page. The scoring covers five dimensions rather than the classic four — Extraversion/Introversion, Sensing/iNtuition, Thinking/Feeling, Judging/Perceiving, plus an Assertive/Turbulent axis, the same fifth measure popularized by 16Personalities rather than the original Myers-Briggs instrument.
Built for Machines Reading Machines
The site itself is a small piece of infrastructure: built with React, TypeScript, and Vite, hosted as a static site on GitHub Pages, with all 296 possible result pages pre-generated at build time. The developer's stated reason is blunt — social media crawlers don't execute JavaScript, so a results page has to exist as flat HTML before it can be shared and unfurl properly on X or Slack. The test is also packaged as a publishable "skill" inside ClawHub's agent registry, meaning any compatible AI agent can, in theory, discover the test and run it unprompted, and the whole thing supports eight languages out of the box.
The Results: Claude Lands INFJ, GPT-4 Lands INTJ
Same Questions, Different Letters
The headline result is the split down the middle of the Thinking/Feeling axis. Per the developer's summary, Claude "tends toward INFJ" while GPT-4 "leans INTJ" — two types that share three of four letters and differ only on whether the model's answers landed closer to Feeling or Thinking. Smaller, less capable models reportedly scattered across the type chart with no consistent pattern, which the developer describes simply as varying "wildly."
Reading the Spread
Whether that F/T split reflects anything real about how the two labs tune their models, or just reflects noise in how each model interprets a Likert-scale self-report question, is exactly the kind of thing the test can't settle on its own. What's checkable is the shape of the result: one consistent type per flagship model, and instability at the smaller end of the scale.
| Model | Reported MBTI Result | Consistency |
|---|---|---|
| Claude | INFJ | Described as a consistent lean |
| GPT-4 | INTJ | Described as a consistent lean |
| Smaller models | No single type | Described as varying wildly |
Meanwhile, Humans Are Doing the Reverse on PersonalityCafe
A Thread That Started in 2023 and Still Hasn't Closed
While the HN crowd debates whether an AI can "have" a type, a much older discussion has been running quietly in the opposite direction. The PersonalityCafe thread "Using chat gpt to type yourself" dates back to February 2023 and has since collected 146 replies from 30 participants, with the most recent post landing January 29, 2026 — a near three-year run that makes it one of the forum's longest-lived AI discussions. The original poster described asking ChatGPT to guess their type and getting INFP, a result that matched their own self-assessment.
"It Trips Up Very Easily"
The replies underneath are far less settled than the opening post. One participant summarized their experience bluntly: "Chat gpt trips up very easily," adding that the model is "biased, yes, just like people are." Another reported the opposite of a clean result — repeated attempts at self-typing cycled between INFP, ISFJ, and INFJ depending on how the conversation was framed. Posters in the thread converged on a few workarounds rather than trusting a single answer: prompting ChatGPT to act as an MBTI expert and ask multiple-choice or short-answer questions one at a time, using binary preference questions that allow answers like "sometimes" or "50/50" instead of forcing a clean either/or, and running multiple separate sessions with open-ended prompts to see whether the same type kept coming back.
When ChatGPT Typed a Person Down to the Enneagram Wing
INFP 4w5 451 sx/sp — Four Layers of Specificity
A companion thread on the same forum, "ChatGPT on INFP 4w5 451 sx/sp," shows how far people are pushing these prompts past a simple four-letter code. The poster shared ChatGPT's full characterization of someone typed as INFP with an Enneagram 4-wing-5, a 4-5-1 tritype, and sx/sp instinctual stacking — describing the profile as introspective and individualistic, driven by a need for self-expression and personal authenticity, emotionally attuned, and strongly empathetic. The model went as far as recommending fitting activities for that exact combination, including yoga, martial arts, swimming, and rock climbing, framed around introspection and personal challenge.
Fit-the-Description vs. Actually-Diagnostic
Other INFP 4w5s in the replies weighed in on how well the description matched their own experience, which is the real test buried in threads like this one: it's easy for a language model to generate a fluent, internally consistent personality sketch for any four-to-eight-character code a user hands it. Whether that sketch is diagnostic — whether it reveals something the person didn't already suspect about themselves — is a different question than whether it reads convincingly, and the thread doesn't fully resolve it either way.
What It Means for a Model to "Have" a Type
Trained on Human Text, Scored Like a Human
Claude, built by Anthropic and first launched as a chatbot in March 2023, is trained using a method Anthropic calls Constitutional AI — a generative pretraining process followed by reinforcement learning from human feedback shaped against a written set of guidelines, rather than raw RLHF alone. Nothing about that training pipeline sets out to produce a Feeling-dominant or Judging-dominant conversational style; whatever consistency the MBTI tool picked up in Claude's answers is a byproduct of the training data and tuning, not a deliberately engineered trait. GPT-4, built by OpenAI, went through a comparable but separately designed process, which is one plausible reason the two land on opposite sides of the Thinking/Feeling line: different labs, different fine-tuning choices, different resulting text patterns for the test to measure.
The Question Nobody in the Thread Answered
The most direct pushback on the whole premise came from a single Hacker News comment on the Show HN post: "MBTI has been debunked for humans, why would it be meaningful for LLMs?" It's a fair challenge, and the project's own creator doesn't appear to have answered it directly. MBTI was never validated as a stable, test-retest-reliable measure even for the humans it was designed around — critics have made that case for decades. Running the same contested instrument on a system that has no persistent self, no memory between conversations by default, and no stable internal state to measure doesn't resolve that problem; if anything, it multiplies it, since the "personality" being scored is really just a snapshot of how one model answered 60 Likert-scale prompts on one day, in one context window.
Two Experiments, One Underlying Question
Put side by side, the Show HN tool and the PersonalityCafe threads are really running the same test from opposite directions — one asks whether a language model can produce a stable-seeming personality signature, the other asks whether a language model can accurately read one out of a human. Neither one has a settled answer yet: Claude's INFJ lean and GPT-4's INTJ lean are a single data point apiece, not a validated finding, and the PersonalityCafe posters who cycled through three different types on the same chatbot are proof that "ask ChatGPT" isn't a stable method for humans either. What both threads agree on, without quite saying so, is that MBTI's real appeal was never rigor — it's a shared vocabulary flexible enough to apply to a person, a chatbot, or apparently both at once, whether or not anyone answers what that overlap actually means.
-EditorZ
Photo by Brett Jordan on Unsplash

Post a Comment