What Is a Situational Judgment Test? The Definitive Guide to SJT Assessment for Executive Hiring and Leadership Development

A situational judgment test (SJT) is a psychometric assessment that evaluates how individuals make decisions and behave when faced with realistic workplace scenarios. Rather than measuring knowledge or technical skills, SJTs assess soft skills, judgment quality, and behavioral tendencies—revealing how a candidate would likely respond to the interpersonal challenges, ethical dilemmas, and leadership decisions they will encounter on the job. These assessments have become central to modern hiring, leadership development, and succession planning because they predict job performance more effectively than many traditional methods and do so at a fraction of the cost of assessment centers or work simulations.

For executives and organizational leaders, SJTs serve a critical strategic function: they identify which candidates possess the judgment, emotional intelligence, and decision-making capability required to lead effectively and navigate organizational risk. This is particularly important in roles where a single misjudgment—whether in conflict resolution, ethical compliance, or stakeholder management—can cascade into significant business consequences. Cernis, a leading provider of executive risk intelligence and psychometric assessment, recognizes SJTs as one of the most evidence-informed, workplace-safe tools for understanding how individuals will behave under pressure, without making diagnostic claims about personality or mental health.

What Is a Situational Judgment Test?

Definition and Core Concept

A situational judgment test presents candidates with a series of realistic workplace scenarios—often written descriptions, video clips, or interactive simulations—and asks them to choose the most appropriate action or rank response options from best to worst. The scenarios are designed to reflect actual challenges that job holders face: managing a conflict between team members, responding to an upset customer, handling an ethical grey area, or making a decision with incomplete information.

The fundamental principle underlying SJTs is straightforward: how a person says they would behave in a hypothetical situation correlates meaningfully with how they actually behave on the job. This is why SJTs are classified as low-fidelity simulations—they simulate the decision-making demands of a role without requiring candidates to physically perform the task (as would happen in an assessment center or work sample). The candidate doesn't actually manage the conflict or serve the customer; they describe how they would handle it. This approach is both practical and valid: it captures judgment without the logistical complexity or cost of high-fidelity simulations.

The term "judgment" in SJT is deliberate. These tests do not have a single objectively correct answer the way a math problem does. Instead, they measure judgment—the quality of decision-making when facing ambiguous, multifaceted situations where different stakeholders have competing needs. A response that demonstrates empathy and collaborative problem-solving might score higher than one that shows quick action but poor communication. The scoring reflects what research and organizational practice have identified as more effective, more professional, or more aligned with the organization's values.

Assessment Method Format Measures Fidelity Cost Scalability
Situational Judgment Test (SJT) Scenarios + response selection/ranking Judgment, soft skills, decision-making, behavioral tendencies Low Low–Moderate High
Assessment Center In-person exercises (role-play, group tasks, presentations) Leadership, teamwork, communication, stress tolerance High High Low–Moderate
Work Sample / Simulation Candidate performs actual job tasks Job-specific skills, task performance High Moderate–High Low–Moderate
Cognitive Ability Test Reasoning, logic, numerical, verbal problems General mental ability, problem-solving capacity N/A Low High
Personality Questionnaire Self-report items (Likert scales) Personality traits, motivations, work style N/A Low High
Structured Interview Standardized questions, behavioral probes Experience, competencies, communication Moderate Moderate Low–Moderate

How SJTs Differ from Other Assessments

While many assessment tools measure knowledge, ability, or stable personality traits, SJTs measure judgment and behavioral choice in context. This distinction is critical for understanding why SJTs have become so popular in hiring and leadership development.

Unlike cognitive ability tests (which measure reasoning capacity), SJTs don't ask "Can you solve this logic problem?" They ask "What would you do if your best performer just told you they're considering leaving?" Unlike personality questionnaires (which measure stable traits), SJTs don't ask "How sociable are you?" They ask "Your team is divided on a key decision. How do you proceed?" Unlike work samples (which measure task execution), SJTs don't ask "Perform this customer service call." They ask "How would you handle an angry customer in this specific situation?"

This focus on judgment in context makes SJTs particularly valuable for predicting job performance. Research shows that SJTs have criterion validity (correlation with actual job performance) comparable to or exceeding that of cognitive tests, and they do so while showing less adverse impact on protected groups. They are also far more cost-effective and scalable than assessment centers. A single assessment center might accommodate 8–12 candidates per day at a cost of $500–$1,500 per person. An SJT can assess hundreds of candidates in a day, online, at a cost of $20–$100 per person. This scalability is why SJTs have become the preferred first-stage screening tool for many large employers.

The Role of SJTs in Executive Assessment

For executive and leadership roles, SJTs take on heightened importance. An executive's judgment—how they navigate ambiguity, manage stakeholder conflicts, make ethical decisions under pressure, and respond to crises—directly affects organizational outcomes and risk. A CFO who misjudges a financial reporting dilemma, a VP of HR who handles a sensitive personnel issue poorly, or a COO who fails to build alignment across functions can create cascading organizational damage.

SJTs designed for executive assessment typically focus on scenarios that reflect C-suite and senior management challenges: managing through organizational change, handling board-level conflicts of interest, responding to reputational threats, making decisions with incomplete or contradictory information, and balancing competing stakeholder interests. These assessments measure managerial judgment—the quality of decision-making when the stakes are high, the information is ambiguous, and the consequences ripple through the organization.

Organizations use executive SJTs for two primary purposes: selection (identifying the strongest external candidates for senior roles) and development (identifying high-potential talent in the pipeline and addressing judgment gaps before they become crises). This dual use is why SJTs are particularly valuable in succession planning and executive risk management. They reveal not just who is ready now, but who has the judgment foundation to be ready in the future.

How Do Situational Judgment Tests Work?

The Anatomy of an SJT Question

Every SJT question follows a consistent structure: a scenario, response options, and a rating or selection task.

The Scenario. The test begins with a realistic workplace situation, typically 50–300 words in length. The scenario provides context: who is involved, what the problem is, what constraints exist, and what makes the situation urgent or important. For example: "You are the director of product development. Your team has just completed a major feature that you believe will be a competitive advantage. However, the VP of Sales has raised concerns that the feature doesn't address the customer feedback from last quarter. Your CEO wants a decision by end of week. Your team is already stretched thin with the next release cycle beginning in three weeks."

Good scenarios are specific enough to evoke a realistic decision-making context but general enough that they don't require specialized job knowledge. They reflect genuine dilemmas where multiple responses could be defensible but some are more effective, more professional, or more aligned with organizational values.

Response Options. Candidates are then presented with 3–6 possible actions they could take. In the product development example above, options might include: "Defer the decision until you can gather more customer data," "Support the Sales VP's concerns and delay the feature," "Proceed with the launch and plan to iterate based on customer feedback," "Schedule a meeting with Sales and Product to align on priorities," or "Ask the CEO to make the decision."

The response options are designed to capture different decision-making styles and competencies. Some options might reflect decisive action, others collaborative problem-solving, others risk aversion, others stakeholder management. The scoring model determines which responses align with effective leadership in that organization.

The Rating Task. Candidates are asked to either (a) select the single best response, (b) rank all responses from best to worst, or (c) rate each response on a scale (e.g., "How appropriate is this action?" on a 1–5 Likert scale). The rating task varies by SJT format and organizational preference. Ranking and rating formats tend to provide more nuanced data than simple selection but require more time from the candidate.

Scoring. Responses are scored by comparing the candidate's choice against a scoring key developed by subject matter experts and validated through research. In simple formats, a correct answer earns 1 point, an incorrect answer earns 0. In more sophisticated scoring, different responses earn partial credit: a highly appropriate response might earn 2 points, a moderately appropriate response 1 point, and an inappropriate response 0 or even negative points. Some SJTs use consensus scoring, where responses are scored based on how a panel of subject matter experts (e.g., senior managers, organizational psychologists) rated them. Others use empirical scoring, where responses are weighted based on which choices correlate most strongly with high job performance among current employees.

Response Formats and Scoring Methods

SJTs come in several response formats, each with different advantages:

Multiple-Choice (Select One Best Response). Candidates choose the single best action from 3–5 options. This is the simplest format, requiring minimal time and cognitive load. Scoring is straightforward: the selected response is either correct or incorrect. However, this format provides less information about how the candidate views other options and may penalize candidates who see merit in multiple approaches.

Ranking (Order All Responses). Candidates rank all response options from most to least appropriate. This format captures more nuance: it reveals not just the candidate's top choice but their overall decision-making hierarchy. Scoring can be done by comparing the candidate's ranking to the expert or empirical key, with partial credit for responses that are in the right general position. Ranking requires more time but yields richer data.

Multiple-Response (Select All That Apply). Candidates select all responses they believe are appropriate, with no ranking. This format works well when multiple actions could be taken simultaneously or in sequence. Scoring counts how many correct responses were selected and how many incorrect responses were included, allowing for partial credit.

Likert Rating Scale. Candidates rate each response on a scale (e.g., 1 = "Highly inappropriate" to 5 = "Highly appropriate"). This format captures the candidate's perception of appropriateness for each option and can reveal fine-grained differences in judgment. Scoring compares the candidate's ratings to the expert or empirical key, often using correlation or distance metrics.

Scoring Models. The scoring key can be developed using three approaches:

Expert Consensus Scoring: A panel of subject matter experts (senior managers, HR leaders, organizational psychologists) independently rates each response option. The consensus rating becomes the scoring key. This approach is practical and defensible but may reflect the panel's biases or assumptions rather than empirical job performance.

Empirical Scoring: The test is administered to current high performers and lower performers in the role. The scoring key is based on which responses correlate most strongly with the high-performer group. This approach is more objective and job-validated but requires a large sample of current employees and may not generalize to different roles or organizations.

Hybrid Scoring: Expert consensus is used to develop the initial key, then empirical data from current employees is used to refine or validate it. This approach combines the practical advantages of expert consensus with the rigor of empirical validation.

Administration and Delivery

Modern SJTs are almost always delivered by computer, either online or at a testing center. Computer delivery allows for easy integration of multimedia (video scenarios, interactive simulations), standardized timing, automatic scoring, and secure proctoring.

Text-Based SJTs. The simplest and most common format uses written scenarios and text response options. Candidates read the scenario and select or rank responses. Text-based SJTs are quick to develop, easy to update, and can be administered on any device with a web browser. Most corporate SJTs (e.g., those used by SHL, Cubiks, and other vendors) are text-based.

Video-Based SJTs. The scenario is presented as a short video (30 seconds to 2 minutes) showing actors in a workplace situation. The candidate watches the video and then selects or rates responses. Video adds realism and emotional context, making the scenario feel more lifelike. However, video SJTs are more expensive to develop (requiring script writing, filming, and editing) and may introduce bias if the actors' appearance, accent, or behavior influences candidate responses. CASPer (Computer-Based Assessment for Situational Judgment), used widely in medical and professional school admissions, uses video scenarios.

Interactive Simulations. The most sophisticated SJTs place candidates in an interactive simulation where they make choices that affect subsequent events. For example, a candidate might choose to "have a one-on-one conversation with the team member" and then be presented with a follow-up video showing how that conversation unfolds, requiring another decision. Interactive SJTs are highly engaging and realistic but are expensive to develop and time-consuming to complete. They are typically used only for high-stakes roles (executive assessment, medical residency selection) or when the organization has the resources to invest in custom development.

Administration Parameters. Most SJTs take 30–90 minutes to complete, depending on the number of scenarios (typically 4–15) and the response format. Some SJTs have no time limit, allowing candidates to work at their own pace. Others impose strict time limits (e.g., 60 minutes for 12 scenarios) to assess decision-making under time pressure. Proctoring may be required (human proctor in a testing center, or AI-based remote proctoring) or may be unproctored (candidates take the test at home without monitoring).

What Do Situational Judgment Tests Measure?

Core Competencies and Soft Skills

SJTs are designed to measure the soft skills and behavioral competencies that predict success in specific roles. Unlike technical skills (which can be taught), soft skills reflect how a person approaches interpersonal and leadership challenges. The competencies measured vary by SJT, but common ones include:

Communication. How clearly and effectively does the candidate express ideas, listen to others, and adapt their message to the audience? Scenarios might involve explaining a difficult decision to a team, delivering bad news to a customer, or negotiating a conflict between departments. High-scoring responses typically demonstrate clarity, empathy, and responsiveness to the other person's perspective.

Teamwork and Collaboration. Does the candidate work well with others, contribute to group goals, and support teammates? Scenarios might involve a team member underperforming, a conflict between colleagues, or a situation where the candidate must rely on others. High-scoring responses typically show a willingness to support others, seek input, and prioritize team success over individual credit.

Leadership and Influence. Can the candidate motivate others, set direction, and influence without formal authority? Scenarios might involve rallying a demoralized team, gaining buy-in for an unpopular decision, or stepping up when a leader is absent. High-scoring responses typically demonstrate vision, confidence, and the ability to inspire and align others.

Problem-Solving and Adaptability. How does the candidate approach novel or ambiguous problems? Do they gather information, consider multiple perspectives, or jump to conclusions? Scenarios might involve a crisis, a change in requirements, or conflicting priorities. High-scoring responses typically show flexibility, creativity, and a willingness to adjust approach based on new information.

Conflict Management. How does the candidate handle disagreement, tension, or opposing viewpoints? Do they avoid conflict, dominate, or seek collaborative solutions? Scenarios might involve a direct report who disagrees with a decision, a customer complaint, or a conflict between peers. High-scoring responses typically demonstrate respect for the other person's perspective, a focus on shared interests, and a commitment to finding solutions that work for all parties.

Ethical Responsibility and Judgment. Does the candidate act with integrity, consider the broader organizational impact, and make decisions aligned with values? Scenarios might involve a temptation to cut corners, a conflict of interest, or pressure to make a decision that violates policy. High-scoring responses typically prioritize ethical principles, transparency, and long-term organizational health over short-term convenience.

Self-Awareness and Emotional Intelligence. Does the candidate recognize their own limitations, seek feedback, and manage their emotions under stress? Scenarios might involve a mistake, criticism, or a situation where emotions are running high. High-scoring responses typically show humility, a willingness to learn, and emotional regulation.

Competency What It Measures Example Scenario Business Impact
Communication Clarity, listening, adaptation to audience Explaining a layoff decision to your team Reduces misunderstanding, builds trust, improves change adoption
Teamwork Collaboration, support, shared goals A colleague is struggling; your deadline is tight Improves retention, reduces silos, increases psychological safety
Leadership Vision, motivation, influence Your team is demoralized after a failed project Improves engagement, accelerates recovery, builds followership
Problem-Solving Flexibility, creativity, information-gathering A key supplier suddenly exits the market Reduces crisis impact, enables innovation, improves agility
Conflict Management Respect, collaborative approach, win-win orientation Two department heads are in a turf war Reduces turnover, improves cross-functional collaboration, prevents escalation
Ethical Judgment Integrity, transparency, long-term thinking Pressure to misrepresent results to meet targets Reduces compliance risk, protects reputation, builds culture of integrity
Emotional Intelligence Self-awareness, emotional regulation, empathy You're criticized publicly in a meeting Improves resilience, reduces interpersonal conflict, enhances executive presence

Behavioral Tendencies and Decision-Making Patterns

Beyond individual competencies, SJTs reveal behavioral patterns and decision-making tendencies. How a candidate responds across multiple scenarios reveals their default approach to workplace challenges. Do they tend to be proactive or reactive? Do they seek input or make unilateral decisions? Do they prioritize relationships or results? Do they take calculated risks or avoid risk?

These patterns are important because they reflect how the candidate will behave in real situations over time. A single scenario response might be luck or a good guess. A pattern across 10 scenarios reflects the candidate's actual decision-making tendency. This is why longer SJTs (with more scenarios) provide more reliable and valid predictions of job performance than shorter ones.

For executive assessment, these patterns are particularly revealing. An executive who consistently chooses collaborative, inclusive decision-making will create a very different culture than one who consistently makes unilateral, decisive choices—even if both approaches can be successful in the right context. Understanding these patterns helps organizations identify candidates whose decision-making style aligns with the organization's values and strategic needs.

Personality and Emotional Intelligence Dimensions

Research has shown that SJTs can be used to measure personality traits and emotional intelligence, particularly dimensions like dependability, conscientiousness, and emotional stability. Unlike personality questionnaires (where candidates self-report their traits), SJTs measure these constructs indirectly through behavioral choice. A candidate who consistently chooses responsible, thoughtful actions across scenarios is demonstrating dependability. A candidate who consistently chooses actions that show emotional regulation and empathy is demonstrating emotional intelligence.

This indirect measurement has advantages: it is less subject to social desirability bias (candidates trying to look good) and it measures traits in context, showing how the person actually applies emotional intelligence in real situations, not just how they perceive themselves.

Where Did Situational Judgment Tests Come From?

Historical Development and Origins

Situational judgment tests emerged in the 1970s and 1980s, initially in military and civil service contexts. The U.S. military developed early SJT-like assessments to evaluate officer candidates' judgment in leadership scenarios. The U.S. Office of Personnel Management (OPM) adopted and formalized the approach, using SJTs to assess federal job applicants. The appeal was clear: SJTs provided a cost-effective, standardized way to measure judgment and soft skills across large numbers of candidates without the expense and logistical complexity of assessment centers.

The term "situational judgment test" itself was formally defined and popularized by organizational psychologists in the 1990s as research demonstrated the validity of the approach. Early academic research (particularly work by Motowidlo and colleagues) established that SJTs had criterion validity—they predicted job performance—and that they were less subject to adverse impact than cognitive ability tests. This research gave organizations confidence that SJTs were both effective and legally defensible.

Evolution and Adoption in Modern Hiring

Throughout the 1990s and 2000s, SJTs migrated from government to corporate hiring. Major assessment vendors like SHL, Cubiks, and Kenexa developed commercial SJT products and integrated them into broader talent assessment suites. By the 2010s, SJTs had become standard practice in corporate recruitment, particularly for graduate hiring, management selection, and customer service roles.

The adoption accelerated further with the rise of video-based SJTs. CASPer (Computer-Based Assessment for Situational Judgment), developed by Altus Assessments in 2010, introduced video scenarios to medical school and professional program admissions. CASPer's success demonstrated that video-based SJTs could be administered at scale to thousands of candidates and could predict success in professional training and practice. Today, CASPer is used by hundreds of medical schools, dental schools, nursing programs, and other professional programs worldwide.

In parallel, interactive and simulation-based SJTs emerged for high-stakes executive and leadership assessment. Organizations began developing custom SJTs tailored to their specific leadership competencies and scenarios. This customization reflected a growing recognition that SJTs are most valid and useful when they reflect the actual challenges and decision-making contexts of the target role and organization.

Research and Validation Milestones

The growth of SJTs has been accompanied by extensive research validating their effectiveness. Key findings include:

Criterion Validity. Meta-analyses have shown that SJTs have moderate-to-strong correlations with job performance (typically r = 0.25 to 0.50, depending on the role and SJT quality). This is comparable to or exceeding the validity of cognitive ability tests and personality assessments, and it is achieved without the adverse impact of cognitive tests.

Predictive Validity. SJTs administered during hiring predict not just overall job performance but specific competencies like leadership effectiveness, teamwork, and customer service quality. This makes SJTs particularly valuable for identifying candidates with high potential in key competency areas.

Incremental Validity. SJTs add predictive value beyond cognitive ability and personality measures. In other words, knowing how a candidate performs on an SJT tells you something about their likely job performance that you wouldn't know from cognitive or personality tests alone. This incremental validity justifies using SJTs in addition to other assessments.

Fairness and Adverse Impact. Research has consistently shown that SJTs have lower adverse impact (differences in average scores between demographic groups) than cognitive ability tests. While SJTs are not bias-free—some research suggests small differences in SJT performance across gender and ethnicity groups—the differences are typically smaller than those found in cognitive testing. This makes SJTs a more legally defensible and equitable assessment method.

Faking and Social Desirability. One concern with SJTs is that candidates might fake their responses to look good. Research on this question shows that while some faking does occur, the degree of faking is less than with personality questionnaires, and importantly, even with faking, SJTs still predict job performance. This suggests that the behavioral patterns revealed by SJTs are robust enough to predict performance even when candidates are trying to present themselves favorably.

Why Do Organizations Use Situational Judgment Tests?

Benefits for Hiring and Talent Selection

Organizations use SJTs in hiring for several compelling reasons:

Cost-Effectiveness. SJTs are inexpensive to administer at scale. A commercial SJT costs $20–$100 per candidate, compared to $500–$1,500 per candidate for an assessment center. For organizations screening hundreds or thousands of candidates, this cost difference is significant. Even custom SJT development (which can cost $20,000–$50,000) pays for itself quickly when used across a large candidate pool.

Scalability. SJTs can be administered to unlimited numbers of candidates simultaneously, online, without geographic constraints. An assessment center can accommodate 8–12 candidates per day. An SJT can assess hundreds in a day. This scalability makes SJTs the ideal first-stage screening tool for high-volume hiring.

Prediction of Job Performance. Research demonstrates that SJTs predict job performance, often as well as or better than interviews, work samples, or personality tests. This predictive validity means that using SJTs in hiring improves the quality of hires and reduces turnover.

Reduction of Bias. While no assessment is completely free of bias, SJTs show lower adverse impact than cognitive ability tests and can be designed to minimize bias through careful scenario development and diverse response options. SJTs also reduce reliance on subjective interviews, which are prone to interviewer bias.

Candidate Experience. Candidates often perceive SJTs as fair and job-relevant. Because the scenarios reflect actual job situations, candidates understand why they are being asked these questions. This perception of fairness improves employer brand and candidate acceptance of hiring decisions.

Leadership Development and Succession Planning

Beyond hiring, organizations use SJTs for leadership development and succession planning. In these contexts, SJTs serve as a diagnostic tool to identify judgment gaps and development needs.

High-Potential Identification. Organizations administer SJTs to high-potential talent in the pipeline to identify who has the judgment and decision-making capability to succeed in senior roles. SJT results, combined with other data (performance ratings, 360 feedback, experience), help organizations identify which high performers are ready for bigger roles and which need additional development.

Succession Planning and Risk Mitigation. In succession planning, organizations need to know not just who is ready now but who has the foundation to be ready in the future. SJTs reveal judgment strengths and gaps, helping organizations plan targeted development. For example, if an identified successor scores well on leadership and strategic thinking but lower on conflict management, the organization can provide coaching or developmental assignments to build that capability before the person steps into the role. This proactive approach reduces succession risk.

Organizational Risk Management. For roles where judgment directly affects organizational risk—executives, board members, risk officers, compliance leaders—SJTs provide evidence about decision-making quality. An executive who consistently demonstrates ethical judgment, stakeholder awareness, and long-term thinking is lower risk than one who shows short-term thinking or ethical grey areas. Organizations increasingly use SJT results as part of executive risk assessment and governance.

Organizational Safety and Ethical Compliance

In regulated industries and roles where judgment directly affects public safety or ethical compliance, SJTs have become essential. Healthcare organizations use SJTs (like CASPer) to assess medical students' and residents' judgment about patient care, professionalism, and ethical dilemmas. Law firms use SJTs to assess candidates' judgment about legal ethics and client relations. Financial services firms use SJTs to assess judgment about compliance and conflicts of interest.

In these contexts, SJTs serve a dual purpose: they identify candidates with strong judgment and they send a signal about what the organization values. A healthcare organization that uses an SJT focused on empathy and patient-centered care is signaling that it values these things. A law firm that uses an SJT focused on ethical judgment is signaling its commitment to professional responsibility.

What Are the Different Types of Situational Judgment Tests?

Response Format Types

As discussed earlier, SJTs vary in how candidates respond to scenarios. The most common formats are:

Choose One (Multiple-Choice). Candidates select the single best response. This is the quickest and simplest format, ideal for high-volume screening. However, it provides less nuanced information than other formats.

Rank All Responses. Candidates order all responses from best to worst. This format captures more information about the candidate's decision-making hierarchy and is more discriminating (it's harder to guess). However, it takes more time and may be frustrating for candidates if they don't see clear distinctions between responses.

Rate Each Response. Candidates rate each response on a scale (e.g., "How appropriate is this action?"). This format captures nuanced judgments and is useful for understanding the candidate's perception of appropriateness for each option. However, it can be time-consuming and may be confusing if the rating scale is not clearly defined.

Select All That Apply. Candidates select all responses they believe are appropriate. This format is useful when multiple actions could be taken and works well for assessing judgment about what actions are acceptable vs. unacceptable. However, it requires careful scoring to handle partial credit.

Scenario Fidelity Levels

Text-Based. The scenario is presented as written text (100–300 words) and responses are text options. This is the most common and practical format. Text-based SJTs are quick to develop, easy to update, and can be administered on any device. They are the default choice for most corporate and academic SJTs.

Video-Based. The scenario is presented as a short video (30 seconds to 2 minutes) showing actors in a workplace or professional situation. Video adds emotional context and realism, making the scenario feel more lifelike. However, video SJTs are more expensive to develop and administer, and there is potential for bias based on actors' appearance, accent, or behavior. Video-based SJTs are commonly used in medical and professional school admissions (e.g., CASPer) and increasingly in corporate assessment.

Interactive/Simulation. The candidate makes a choice that affects what happens next, creating a branching scenario. For example, "You choose to have a one-on-one conversation with the team member" and then a video shows how that conversation unfolds, presenting a new decision point. Interactive SJTs are highly engaging and realistic but are expensive to develop and time-consuming to complete. They are typically used only for high-stakes selection (executive roles, medical residency) or when the organization has the resources to invest.

Domain-Specific SJTs

SJTs are increasingly customized to specific domains and roles. Some examples:

Medical SJTs (CASPer, MCAT SJT). Focus on scenarios involving patient care, professionalism, ethical dilemmas, teamwork in healthcare settings. Used in medical school admissions and residency selection.

Legal SJTs. Focus on scenarios involving client relations, ethical judgment, professional responsibility, conflicts of interest. Used in law school admissions and law firm hiring.

Management/Leadership SJTs. Focus on scenarios involving team management, conflict resolution, decision-making, strategic thinking, change management. Used in corporate hiring for management roles and executive assessment.

Customer Service SJTs. Focus on scenarios involving difficult customers, service recovery, communication, empathy. Used in hiring for customer-facing roles.

Sales SJTs. Focus on scenarios involving client objections, negotiation, relationship building, ethical boundaries. Used in hiring for sales roles.

How Valid and Reliable Are Situational Judgment Tests?

Criterion Validity and Predictive Power

Criterion validity is the correlation between test scores and actual job performance. Research on SJTs shows moderate-to-strong criterion validity: meta-analyses typically report correlations of 0.25 to 0.50 with job performance ratings. This is comparable to or exceeding the validity of cognitive ability tests (which average around 0.25 to 0.40 for many jobs) and personality assessments (which average around 0.15 to 0.30).

The validity of SJTs varies by several factors: (1) the quality of SJT development (well-developed SJTs with expert consensus or empirical scoring are more valid), (2) the job type (SJTs tend to be more valid for jobs with high interpersonal demands), (3) the similarity between the SJT scenarios and actual job situations (high similarity = higher validity), and (4) the scoring method (empirically derived keys tend to be more valid than expert consensus keys).

For executive and leadership roles, SJTs tend to show particularly strong validity because the soft skills and judgment they measure—leadership, communication, conflict management, strategic thinking—are directly related to executive performance. An executive's ability to handle ambiguity, make decisions with incomplete information, and navigate stakeholder conflicts is central to their effectiveness, and SJTs measure exactly these capabilities.

Reliability and Consistency

Reliability refers to the consistency of test scores. An SJT is reliable if a candidate would receive similar scores if they took the test multiple times (test-retest reliability) and if different raters score the same responses similarly (scorer reliability).

Internal Consistency. SJTs typically show good internal consistency (Cronbach's alpha of 0.60 to 0.80), meaning that responses across different scenarios correlate with each other. This suggests that the SJT is measuring a coherent construct (e.g., judgment quality) across scenarios.

Test-Retest Reliability. Research on test-retest reliability (administering the same SJT twice to the same candidates) shows moderate stability (correlations of 0.40 to 0.70), with some variation depending on the time interval and scenario similarity. This moderate stability is expected: some variation in responses is normal and reflects the candidate's genuine variability in judgment across different scenarios. However, the overall pattern of judgment (e.g., whether someone is generally collaborative vs. directive) tends to be stable.

Scorer Reliability. For SJTs with expert consensus or empirical scoring, scorer reliability is typically high (>0.90) because the scoring key is clearly defined. For SJTs that require subjective judgment in scoring (e.g., open-ended responses), scorer reliability can be lower and requires careful training of raters.

Fairness and Bias Considerations

A critical question for any assessment is whether it is fair and unbiased. Research on SJT fairness has found:

Adverse Impact. SJTs typically show lower adverse impact (differences in average scores between demographic groups) than cognitive ability tests. Meta-analyses show that while there are sometimes small differences in SJT scores by gender and ethnicity, these differences are typically smaller than differences in cognitive test scores. This makes SJTs a more equitable assessment method and more legally defensible under employment law.

Cultural Fairness. Because SJTs measure judgment in context rather than abstract reasoning, they may be less subject to cultural bias than cognitive tests. However, the scenarios themselves must be culturally fair and relevant. An SJT with scenarios that only reflect one cultural or organizational context may be biased against candidates from other backgrounds. Well-designed SJTs include diverse scenarios and response options that reflect multiple perspectives and cultural contexts.

Stereotype Threat. Stereotype threat—the anxiety that arises when a person is aware of negative stereotypes about their group's performance on a task—can affect test performance. Some research suggests that SJTs may be less subject to stereotype threat than cognitive tests, possibly because the scenarios are less obviously "ability-based." However, this research is still developing.

Coaching and Preparation. One fairness concern is whether candidates can improve their SJT scores through coaching or practice. Research shows that while some improvement occurs with practice, the gains are typically modest (10–20% improvement) and less than the gains possible with cognitive tests. This suggests that SJTs are relatively resistant to coaching and that differences in SJT scores reflect genuine differences in judgment rather than test-taking skills.

Frequently Asked Questions About Situational Judgment Tests

What is the difference between a situational judgment test and an assessment center?

Assessment centers and SJTs both measure soft skills and judgment, but they differ in format and fidelity. Assessment centers place candidates in a physical setting and ask them to perform tasks (role-play a difficult conversation, lead a group discussion, solve a business problem) while trained assessors observe. SJTs ask candidates to describe how they would handle hypothetical situations. Assessment centers are high-fidelity (candidates actually perform tasks) and intensive (typically 6–8 hours, 1–2 days), while SJTs are low-fidelity (candidates describe responses) and quick (30–90 minutes). Assessment centers are more expensive ($500–$1,500 per person) and less scalable (8–12 candidates per day) but may be more realistic and engaging. SJTs are cheaper and more scalable but provide less behavioral observation. Many organizations use SJTs as a first-stage screening tool and assessment centers for final-stage selection of top candidates.

Can you prepare for a situational judgment test?

Yes, candidates can prepare for SJTs through practice and study, but the gains are typically modest. Preparation strategies include: (1) understanding the competencies being measured, (2) reviewing example scenarios and thinking about how you would respond, (3) considering different perspectives and stakeholder interests in each scenario, (4) practicing with sample tests, and (5) getting feedback on your responses. However, because SJTs measure judgment and decision-making tendencies, extensive practice or coaching cannot dramatically change your scores. Research shows that practice effects are smaller for SJTs than for cognitive tests, suggesting that SJT scores reflect genuine differences in judgment.

How long does a typical situational judgment test take?

Most SJTs take 30–90 minutes to complete, depending on the number of scenarios (typically 4–15) and the response format (select one, rank all, or rate each). Some SJTs have no time limit, allowing candidates to work at their own pace. Others impose strict time limits to assess decision-making under time pressure. Video-based SJTs may take longer because candidates must watch video scenarios. Organizations should provide clear time guidance to candidates before the test begins.

What is a good score on a situational judgment test?

SJT scores are typically reported as a percentage correct (e.g., "75% correct") or as a percentile rank (e.g., "75th percentile compared to other candidates"). What constitutes a "good" score depends on the organization's hiring standards and the competitiveness of the candidate pool. For high-stakes selection (executive roles, professional school admissions), organizations often use a cutoff score (e.g., "candidates must score in the top 25%") to narrow the candidate pool. For development purposes, scores are less about meeting a threshold and more about identifying strengths and development areas. Organizations should establish scoring benchmarks based on current high performers in the role.

Are situational judgment tests used for executive assessment?

Yes, absolutely. SJTs are increasingly used for executive assessment, both for external hiring and for internal succession planning and development. Executive-level SJTs focus on scenarios involving strategic decision-making, stakeholder management, crisis response, ethical judgment, and organizational change. Organizations use executive SJTs to identify candidates with strong judgment for C-suite and senior management roles and to identify high-potential talent in the pipeline who need development before stepping into executive positions.

How do organizations use SJT results?

Organizations use SJT results in several ways: (1) Hiring screening: SJTs are used as a first-stage screening tool to narrow the candidate pool before interviews or assessment centers. (2) Hiring decision: SJT scores are one input (along with interviews, references, other assessments) into final hiring decisions. (3) Development: SJT results are used to identify judgment strengths and development areas, informing coaching, training, or developmental assignments. (4) Succession planning: SJT results are combined with other data to identify high-potential talent and assess readiness for higher roles. (5) Organizational risk assessment: For sensitive roles, SJT results are used to assess judgment quality and decision-making risk.

What does research say about SJT validity?

Research on SJTs has been extensive and consistently positive. Meta-analyses show that SJTs have criterion validity (correlation with job performance) of 0.25 to 0.50, which is comparable to or exceeding that of cognitive ability tests and personality assessments. SJTs also show incremental validity—they predict job performance beyond what cognitive and personality tests predict. SJTs show lower adverse impact than cognitive tests, making them more equitable. Research also supports the construct validity of SJTs—they measure what they claim to measure (judgment, soft skills, behavioral tendencies)—and shows that SJT scores are relatively stable over time, suggesting they measure enduring aspects of judgment.

Can SJTs be biased?

No assessment is completely free of bias, but SJTs show lower bias than many alternatives. Research shows that SJTs have lower adverse impact than cognitive ability tests and can be designed to minimize bias through careful scenario development and diverse response options. However, bias can occur if: (1) scenarios are culturally specific or reflect only one perspective, (2) response options are not equally accessible to candidates from different backgrounds, or (3) the scoring key reflects the biases of the panel of experts who developed it. To minimize bias, organizations should use diverse scenario development teams, include scenarios relevant to diverse candidates, and validate scoring keys empirically across demographic groups.

How much does a situational judgment test cost?

The cost of SJTs varies widely depending on whether you use a commercial off-the-shelf test or develop a custom test. Commercial SJTs: $20–$100 per candidate (one-time licensing fee plus per-candidate administration cost). Custom SJT development: $20,000–$50,000+ depending on complexity (text-based vs. video-based), number of scenarios, and level of customization. For organizations with high-volume hiring, the per-candidate cost of custom SJTs can be competitive with commercial tests once you amortize development costs across a large candidate pool.

What is the difference between CASPer and other SJTs?

CASPer (Computer-Based Assessment for Situational Judgment) is a video-based SJT developed by Altus Assessments for professional school admissions, particularly medical, dental, and nursing schools. CASPer presents video scenarios (actors in realistic professional situations) and asks candidates to type written responses. CASPer is unique in several ways: (1) it uses video rather than text scenarios, adding realism and emotional context, (2) it uses open-ended written responses rather than multiple-choice, requiring candidates to articulate their reasoning, and (3) it is used primarily in academic admissions rather than corporate hiring. Other commercial SJTs (e.g., SHL, Cubiks, Kenexa) typically use text scenarios and multiple-choice or ranking responses and are used primarily in corporate hiring. Both CASPer and traditional SJTs are valid predictors of professional performance, but they differ in format and context.