Skip to content

Building a question bank that lasts: tagging, difficulty and reuse

Every training institute has the same file somewhere: Questions_Final_v3_USE_THIS.docx. Nobody remembers which questions came from the 2023 paper, which ones half the batch got wrong for no good reason, and which ones are circulating on a group chat. So each term someone writes fresh questions, and the cycle repeats.

A question bank is the alternative. Not a folder of papers, but a set of individually tagged, individually measured questions you draw papers from. Built properly, it improves every time you use it, because each sitting tells you which questions are working.

This is how to build one that lasts: tagging by topic and outcome, difficulty you measure rather than guess, stems and distractors that hold up, a review workflow two people can run, retiring leaked questions, and reusing the bank across batches.

Tag every question the day you write it

Tagging later never happens. The person who wrote the question is the only one who knows what it was for, and by next term they have forgotten. Decide your tag fields once and make them compulsory at entry.

Field Example What it buys you
Topic GST / Input tax credit Draws a balanced paper
Learning outcome Can identify blocked credits Tells you what a low score means
Cognitive level Recall, apply, analyse Stops a paper being all recall
Type MCQ, short answer, descriptive Controls grading effort
Difficulty Easy, medium, hard Keeps papers comparable across batches
Status Draft, reviewed, live, rested, retired Stops an unreviewed question reaching a learner
Source and author Batch 12 internal, A. Kulkarni Lets you trace a leak or a bad key
Last used June 2026 Prevents the same question two terms running

The outcome tag is worth arguing about. Topic tells you a learner is weak in a chapter. Outcome tells you what they cannot do, which is the only one useful to a teacher on Monday morning.

Difficulty you measure, not difficulty you guess

Authors are poor judges of difficulty, because they know the answer. After each sitting, let the data relabel your questions.

Two standard measures do the work. The difficulty index is the correct responses from the high scoring group plus those from the low scoring group, divided by the total number of responses. An index near 50 per cent sits in the middle: higher suggests the item may be too easy, lower that it may be too hard.

The discrimination index compares the two groups: correct responses from the high group minus those from the low group, divided by half the total responses. It runs from minus 1.0 to plus 1.0. A positive value means more of your strong learners got it right, which is what you want. A negative value means more of your weak learners got it right, which almost always points to a wrong key, an ambiguous stem or a distractor that is accidentally defensible.

The routine after every exam is short. Sort by discrimination index, fix or retire everything negative or near zero, then look at the difficulty extremes: items nobody got right are usually broken rather than hard.

Writing stems and distractors that hold up

The NBME item writing guide is the most useful free reference here, and its central test is the cover the options rule: a test taker should usually be able to read the vignette and lead-in, cover the options, and guess the correct answer without seeing the option set. If they cannot, the options are doing work the stem should have done.

Work through this checklist before a question goes live.

  • The lead-in is focused, closed and clear, and answerable from the stem alone.
  • Four or more options, all plausible to someone who has not learnt the material.
  • Options are grammatically parallel and follow from the lead-in.
  • No “none of the above”; name the specific action instead.
  • No negatively phrased lead-in using EXCEPT, which test takers miss.
  • No absolute terms such as always or never in the options.
  • No word repeated between the stem and the correct answer.
  • The correct answer is not the longest or most detailed option.
  • No convergence, where the correct answer is the option with the most in common with the others.

Distractors are where most banks are weak. A good distractor is a mistake a real learner makes: the wrong formula, the common misreading, last year’s rule. Options invented to fill space are why an item is answered correctly by almost everyone and tells you nothing.

A review workflow two people can run

The guide is blunt about why review matters: each item should be reviewed to identify and remove technical flaws that add irrelevant difficulty or benefit savvy test takers. That does not need a committee. It needs one author and one reviewer who is not the author.

  1. Author writes the question, sets all tags, writes the key and a one line justification for it.
  2. Reviewer applies the checklist above, answers the question cold without looking at the key, and either approves it or sends it back with a reason.
  3. Status moves to live. Only live questions are eligible to be drawn into a paper.
  4. After the sitting, whoever owns the bank reviews the item statistics and sets the status: keep, revise, rest or retire.

Two rules make this stick. Nobody reviews their own question, and a question with no reviewer name against it cannot go live.

If you are drafting a large batch of questions from a syllabus or a PDF, treat the output as a first draft entering this workflow at step one, never as finished items. The same discipline applies to any drafted course content a human has to review before learners see it.

Retiring leaked and worn out questions

Questions wear out. Once enough learners have seen an item, it stops measuring learning and starts measuring who has the old paper. Leakage of a question paper or answer key, or part of it, is listed as an unfair means under the Public Examinations (Prevention of Unfair Means) Act, 2024, which covers notified public examinations. Your internal tests sit outside that Act, but the failure is the same.

Signs an item has leaked or worn out:

  • Difficulty index jumps sharply between two sittings with no change to the teaching.
  • Average answer time falls well below the rest of the paper.
  • Discrimination index collapses towards zero while the score stays high.
  • The item appears verbatim in a shared document or group.

Retire, do not delete. Mark the item retired with a date and a reason so nobody re-adds it next year, and keep a rested status for items you want back after two or three terms out of circulation. Make sure the bank is deep enough that retiring twenty questions does not leave a topic bare: a working rule is three to four times as many live questions per topic as any single paper draws from it.

Reusing the question bank across batches

The payoff comes when the bank outlives the batch. Define papers as recipes against tags rather than as fixed lists, so the same recipe produces a fresh paper each term: six questions from Topic A at medium difficulty, three from Topic B at hard, two applied items from Topic C.

Then keep the bank separate from the course. Courses get restructured, renamed and retired; the questions should not go with them. Tagging by outcome rather than chapter number makes that possible, because outcomes survive a syllabus revision.

Finally, report at question level, not just paper level. A score of 62 per cent tells a learner nothing they can act on. A breakdown showing they are solid on two outcomes and weak on a third is a study plan, which is the case our note on an assessment engine that measures learning makes at length.

Frequently asked questions

How many questions does a bank need?

Enough that no learner sees the same paper twice and no topic runs bare after retirements. Count backwards from the paper: if a topic supplies eight questions and you run three sittings a year, thirty to forty live questions on that topic is a workable floor. Depth matters more than total size.

Do I need to rewrite questions for every batch?

No. That is the point of measuring items. A question with a good discrimination index and a sensible difficulty index is an asset; reusing it across batches is what makes results comparable. Rewrite only what the statistics or a leak tells you to.

Can descriptive questions live in the same bank?

Yes, with the same tags plus a marking scheme attached to the question rather than the paper. You lose auto grading and item statistics, so keep descriptive items for outcomes that genuinely need them and let objective items carry the measurement load.

Keeping the bank in one place

Quipu LMS holds question banks for MCQ, short answer and descriptive items alongside an assessment engine that randomises, times and auto grades, so the tags you set and the statistics each sitting produces stay attached to the same questions.

Sources

Keep reading

WhatsApp