Exampractice
Exam Preparation

How Certification Exam Difficulty Is Determined

Cut scores, scaled scoring, item calibration and adaptive testing — how certification programmes engineer and measure exam difficulty behind the scenes.

Amara Okafor · 7 min read
Conveyor-belt diagram showing how an exam question moves from writing through piloting and calibration to scaled scoring.

Here is a puzzle that confuses almost every first-time candidate: CompTIA Security+ requires 750 on a scale that runs from 100 to 900, AWS Cloud Practitioner requires 700 out of 1,000, and neither number is a percentage of questions answered correctly. If passing scores are not percentages, what are they — and who decided that this level of performance counts as competent?

Short answer: certification difficulty is engineered, not accidental. Programmes decide what a minimally competent professional must know, set a passing standard (the cut score) through structured expert judgement, trial every question on real candidates before it counts, and report results on a scaled score so that different versions of the exam are equally hard to pass. This article walks through that machinery stage by stage.

One boundary note: this piece explains the objective side of difficulty — the measurement and scoring apparatus. The design levers that make questions demanding in the first place are covered in what makes a certification exam difficult, and the human factors that make an exam feel harder than it measurably is belong to why some certification exams feel harder than others.

Stage one: defining the minimally competent candidate

Every defensible exam starts with a job-task analysis: the programme surveys practitioners and subject-matter experts to establish what someone in the target role actually does, then converts that into the published exam objectives and their domain weightings. This is why blueprints change every few years — the job changes, so the exam must follow.

Crucially, the reference point for the whole system is the minimally competent candidate: not an expert, not a star, but the borderline professional who just clears the bar for safe, effective practice. Every later decision — how hard questions should be, where the pass mark sits — is anchored to this hypothetical person. Difficulty, in psychometric terms, is always difficulty relative to minimal competence, which is why exams aimed at experienced professionals can legitimately feel severe to newcomers without being unfair.

Stage two: setting the cut score

The passing standard is called a cut score, and reputable programmes do not pluck it from the air. The most widely used family of standard-setting approaches is associated with the Angoff method, in which a panel of subject-matter experts examines each question and estimates the probability that a minimally competent candidate would answer it correctly. Aggregate those judgements across the whole question pool and you get a defensible passing standard rooted in expert consensus rather than tradition or round numbers.

Two consequences follow that candidates rarely appreciate:

  • The cut score reflects the questions, not a fixed percentage. If the expert panel judges the item pool to be hard, the resulting standard corresponds to fewer raw correct answers; if the pool is gentler, more. The standard tracks competence, not an arbitrary "70%".
  • There is no grading curve in the classroom sense. Certification exams are criterion-referenced: you are measured against the standard, not against the other people testing that day. In principle, everyone in the room can pass.

Providers publish the outcome of this process as their passing score — CompTIA lists 675 out of 900 for A+ Core 1 and 700 out of 900 for Core 2; AWS lists 720 out of 1,000 for Solutions Architect – Associate — but the standard-setting deliberations themselves stay confidential. For exact methodology on a specific exam, the provider's certification handbook or exam guide is the only authoritative source.

Stage three: calibrating questions on real candidates

A question's true difficulty cannot be known from reading it — it has to be measured. So programmes pilot new items by embedding them, unscored and unmarked, inside live exams. AWS is unusually transparent about this: its exam guides state that 15 of the 65 questions on its Cloud Practitioner and Solutions Architect – Associate exams are unscored items that "do not affect your score", included to gather performance data, and candidates are not told which they are.

Piloting yields statistics for every item: what proportion of candidates answer it correctly, and how well it separates strong candidates from weak ones. Items that prove ambiguous, too easy, or misleading are revised or discarded before they ever count. Modern programmes typically analyse this data within frameworks psychometricians group under item response theory — models that estimate each item's difficulty and discriminating power on a common scale, allowing questions to be compared and assembled into balanced forms. The mathematics varies by programme and is rarely published; what matters for candidates is the consequence: by the time a question counts against you, its difficulty is a measured property, not a guess.

This is also worth internalising for exam day: since pilot items are invisible, a bafflingly obscure question may simply be an experiment that will never affect your result. Answer it and move on.

Stage four: scaled scoring — making different exams equally hard

No two candidates necessarily see the same set of questions. Programmes maintain large item pools and assemble multiple exam forms, which inevitably differ slightly in average difficulty. Reporting raw percentages would therefore be unfair: 80% on a harder form represents more ability than 80% on an easier one.

The fix is the scaled score. Raw performance is converted onto a fixed reporting scale — 100–900 for CompTIA, 100–1,000 for AWS — through a statistical equating process that compensates for form differences. AWS's exam guides say this explicitly: scaled scoring exists to equate results across exam forms of slightly varying difficulty. A 750 on one Security+ sitting and a 750 on another represent the same estimated ability, even if the two candidates answered different numbers of questions correctly.

This resolves the opening puzzle. The passing score is not 83% or 75% of anything; it is a point on an equated ability scale. It also explains a common candidate complaint — "my score report seems inconsistent with how many I think I got wrong" — since the raw-to-scaled conversion differs by form and is never published.

Three further scoring rules shape effective strategy, all verified in AWS's guides and worth checking with any provider:

  1. Unanswered questions score as incorrect, and there is no penalty for guessing — so blanks are pure loss.
  2. Scoring is compensatory: you need to pass the exam overall, not every domain individually. A weak domain can be carried by strong ones.
  3. Unscored pilots are invisible, so perceived performance and reported score can legitimately diverge.

A variant: adaptive testing

Some programmes go a step further and let difficulty respond to you in real time. In computerised adaptive testing, the engine selects each successive question partly based on your performance so far — answer well and the questions get harder, because harder questions extract more information about high-ability candidates. Where a programme uses an adaptive format, the practical upshot is counterintuitive: an exam that feels punishingly difficult may mean the engine rates you highly. Whether a given certification uses adaptive delivery, and under what rules, varies by programme — check the provider's exam-format documentation rather than assuming, as most mainstream IT certification exams remain fixed-form.

What all this machinery means for your preparation

Understanding the apparatus changes how you should interpret your own results:

  • Stop converting practice percentages into predicted scaled scores. No public formula connects them. Use practice results diagnostically — which domains are weak, which question types slow you down — rather than as a score forecast. A timed practice test simulation is most useful as a domain-by-domain readiness check under realistic conditions, and how to improve your certification exam score covers reading a real score report after the fact.
  • Trust the blueprint weightings. They come from the job-task analysis and govern how forms are assembled; your study time should mirror them.
  • Never leave blanks where your provider confirms there is no guessing penalty.
  • Do not fear the borderline. The cut score was set so that minimal competence passes. You are not chasing perfection; you are demonstrating that you clear a carefully engineered bar.

The bar is engineered — study to its specification

A certification exam's difficulty is the output of a deliberate pipeline: a job analysis defines competence, an expert panel converts it into a cut score, live piloting measures every question, and scaled scoring holds the standard steady across forms and years. None of it is arbitrary, and almost all of it is documented — the passing scale, the scored/unscored split where disclosed, and the domain weightings sit in the official exam guide of whichever certification you are pursuing. Read that document as the specification it is, then browse the certification exams directory to find preparation materials matched to the exact exam code you will face.

Frequently asked questions

Do certification bodies publish their pass rates or cut-score methods?

Generally no. Most major providers treat pass rates and standard-setting records as confidential. What they do publish — passing scores, scales, formats, domain weightings — appears in official exam guides, which is why those documents outrank any third-party difficulty ranking.

Can the passing score change over time?

Yes. When a programme releases a new exam version after a fresh job-task analysis, it re-runs standard-setting, and the published passing score can move. Always check the current exam guide for the exact code you are booking rather than relying on figures for a predecessor version.

If scoring is compensatory, can I skip studying my weakest domain?

Risky. Compensatory scoring (confirmed by AWS for its exams) means a weak domain will not automatically fail you, but domains are sampled in proportion to their published weightings — a heavily weighted blank spot can sink the total on its own.

Exam facts in this guide were checked against official certification-provider pages on . Fees, exam codes and policies change — confirm on the provider’s own site before you book.

Put it into practice

Test what you have just read

Reading about an exam only takes you so far. Work through practice questions for your certification and find the gaps before exam day does.

You may also like