
When AI writes multiple-choice questions at scale, sometimes the answer is too obvious. Removing the question stem helped us figure out why.
This study is part of a larger WGU Labs project called Current Skills Validation (CSV). CSV is an AI system that recognizes the skills people build outside traditional schooling, such as through work, caregiving, military service, and self-directed learning. The system generates its own assessments rather than drawing on a fixed, hand-written question bank.
The scale of generating assessments in real-time is what makes CSV useful, and it's also what makes quality assurance important: if we want the system to produce trustworthy evidence of what a learner actually knows, we have to be able to check the quality of the questions it tends to generate. Multiple-choice questions (MCQs) are one of the question types CSV generates. This study takes on a specific thing that can undermine an MCQ's quality: guessability, or an individual’s ability to pick the right answer even if they don’t know anything about the topic.
We’re still building the CSV system. This article highlights how we benchmarked the system and where we’ll focus for the next iteration.
Why guessability matters
For multiple-choice questions, we assume that the chance of guessing the right answer is 100% divided by the number of answer options. If the options are well-constructed, someone who knows nothing about the topic should score at chance: on a 3-option question, a learner guessing at random should score ~33%. Many models of knowledge embedded in intelligent tutoring systems or used to estimate the difficulty of tests use chance as a meaningful factor, so it matters that the chance we assume on paper matches the chance we actually see when real students sit down with the questions.
Second, guessability affects whether we are measuring what we intend to measure. If someone who doesn’t know the topic can pick out the right answer, that is a signal to us that the question isn’t measuring what we want it to. It may instead be measuring English language ability, or “test-taking savvy.” We call this construct-irrelevance variance, and it weakens any claim we want to make about what a score means.
Third, guessability is a matter of equity. Students from wealthier school districts often get more practice taking high-stakes tests and more explicit instruction in test-taking strategies. If a question is not carefully built, it may favor students who know those strategies over students who don't, even when the two know the topic equally well.
The equity aspect is what brought us to this work. We wanted to see whether students who didn't know much about a topic could still figure out the right answer to multiple-choice questions produced by the AI system we're developing. If students could figure out the right answer far more often than chance, that would tell whether the question itself was giving away the answer.
The two places guessability comes from
When you take a multiple-choice question apart, there are two main drivers of guessability.
The first is the question stem, which sets up what's being asked. For example, "In nursing, what is the purpose of a comprehensive assessment?" A stem can be transparent enough that a test-taker knows the answer before seeing any options. The familiar way to probe this is to present the stem alone, like a short-answer question, and ask the test-taker to fill in the answer.
The second is the answer options, which include the correct answer plus the incorrect ones, which we call alternatives. Options can give the answer away on their own.
Here's an example:
During recess, a third-grade student tells you, "My stepfather hit me with a belt last night." The student is calm and asks you not to tell anyone. What should you do first?
A. Promise to keep the report private until the student is ready to talk more.
B. Report the disclosure immediately through your school's mandated reporting process. [Correct]
C. Ask the student detailed questions and confirm the injury before you report it.
The correct answer sounds more professional and runs longer than the alternatives. That alone clues a test-taker into which one to pick.
The two components can also work together. Look at this example:
Under FERPA, a student’s birth date is an example of an:
A. Directory information
B. Education record [Correct]
C. Confidential file
The wrong answers don't line up grammatically with the stem. The "an" signals that the next word starts with a vowel, which is only true for the correct answer. Thus, an experienced test-taker can figure out the right answer without knowing any of the content.
Figuring out which component is to blame
When you're improving questions, you want to know which part to spend time on. That's a question of causal inference: isolating the independent effect of one thing. Presenting the whole question tells you whether it's guessable overall, but not why.
Based on our internal testing, we suspected that the answer options our question generator produced were driving guessability, so we developed two hypotheses:
- If the questions were guessable, then participants who didn’t know much about the topic would score well above chance, indicating that they are doing something other than random guessing.
- If participants scored about the same on complete questions and questions with the question stem removed, then guessability was being driven by the answer options.
What we did
Our goal was a quick proof-of-concept to inform rapid iteration. That meant we selected a small sample size and didn't need statistically significant results. As we move along in the development process, we plan on rerunning this benchmarking procedure with increasingly larger samples to offer stronger claims about the CSV system.
We recruited 12 undergraduate students from the Student Insights Council at Western Governors University. All were enrolled in technology programs (e.g., computer science). Students rated their knowledge of elementary education and nursing on a scale from 1 (not knowledgeable at all) to 10 (extremely knowledgeable). Students tended to not know much about either topic: Elementary Education: M = 4, range = 1 - 8, mode = 1; Nursing: M = 3, range = 1 - 6, mode = 1, which is what we were hoping for.
Each participant was randomly assigned to answer 20 questions in either education or nursing. Within those 20, half were presented complete and half with the stem removed. We counterbalanced whether participants saw the complete or no-stem questions first, and we randomized the order of questions and of options within each question, as a precaution against order effects. After the study, we asked students if they looked up the answers to any question, and no student reported doing so. Here's each question form.
Complete question: An elementary teacher checks a student’s weekly reading records and notices that accuracy is high, but the student reads very slowly and loses meaning when asked to retell. What does this pattern most likely mean?
- The student has strong overall reading comprehension.
- The student is decoding accurately but may not yet be reading fluently enough for meaning.
- The student is decoding inaccurately and needs basic phonics review.
No-stem question: Even though we are not giving you the question, pick what you think the correct answer is.
- The student has strong overall reading comprehension.
- The student is decoding accurately but may not yet be reading fluently enough for meaning.
- The student is decoding inaccurately and needs basic phonics review.
What we found: Both hypotheses held

We focused on means rather than statistical significance. Even so, the results are striking. Both hypotheses held. Learners who didn't know much about the topics scored well above chance, which suggests they were doing more than guessing at random. And there wasn't much difference between the complete questions and the no-stem questions. That small gap points to the answer options as the source of guessability: the alternatives probably weren't competitive enough, so the right answer jumped out.
Our next question was why the right answer jumped out. To get at that, we did a qualitative review of the questions students guessed especially well on. Two examples showcase trends we found.
Example 1: Which term describes believing nursing ability can improve through effort, feedback, and practice?
A. Fixed mindset
B. Growth mindset [Correct]
C. Reflective practice
At first glance this looks hard to guess, as reflective practice (i.e., the process of analyzing your actions, experiences, and decisions to learn and improve) is pretty close. What gives the answer away is "fixed mindset," which is the opposite of the correct answer. The two share the keyword "mindset" and form an obvious pair, so a test-taker can infer that one member of the paired term is the intended answer without knowing the concept.
Example 2: A fourth-grade student with a Section 504 plan has an accommodation for extended time on classroom and state assessments. During a timed math quiz, the student asks for the extra time listed on the plan. What should the teacher do?
A. Wait and offer the extra time only if the student appears to be struggling
B. Provide the extended time exactly as listed in the 504 plan [Correct]
C. Adjust the time informally with the student based on how hard the quiz seems
This was another highly guessable question. When we reviewed it, the issue was that only the correct answer used a professional term, the "504 plan," while the other options sounded informal, even if they address a plausible misconception (i.e., only give the student extra time if they seem like they need it in the moment).
The takeaway from these questions is that an alternative can look plausible and competitive on its own, but the construction of the surrounding set changes how a test-taker reads it. That kind of guessability only shows up when you look at the option set as a unit. These results have led us to think about guidelines for our system that work on sets of options rather than individual alternatives.
Why this matters for how we build multiple-choice questions
With this study, we have a proof-of-concept for a new method: testing guessability by presenting only the answer options. Doing that let us make a more precise inference that guessability was coming from the options themselves rather than from the stem or the interaction between stem and options. The method needs no special instrumentation and transfers to any multiple-choice bank, including AI-generated ones.
The stronger causal inference from this approach is key, as it tells us which component to iterate on instead of revising blindly. Addressing guessability also helps our system move towards equity for students with different levels of test-taking strategy knowledge. We’ll keep using this approach as we benchmark future iterations of the system.

