Group of Architects working in modern start up office

Personality assessments have become a fixture of modern management. Teams use them to onboard new members, coach existing employees, build stronger working relationships, and make sense of why colleagues who are equally talented can approach the same problem in completely different ways. DISC is one of the most widely used frameworks for this kind of work, chiefly because it focuses on observable workplace behaviors. It’s easy to apply on a Monday morning without a psychology degree.

Unfortunately, that accessibility is also DISC's biggest vulnerability. Because the model is simple to explain, it is also simple to imitate. A search for “free DISC test” turns up dozens of options, many of which were built by someone with an eye for design and a weekend rather than a research team and a validation study. For a manager deciding how to invest a training budget, or an HR leader choosing a tool that will shape team building and development decisions, the difference between a rigorously built assessment and a plausible-looking imitation determines whether the results mean anything at all.

So what actually separates a scientifically credible DISC assessment from one that simply looks the part? The answer comes down to two things: 

  • Reliability – how consistently it measures what it claims to measure, and
  • Validity – how well its results line up with real behavior at work.

A Quick Refresher on DISC

DISC traces back to the 1920s and the work of psychologist William Moulton Marston, who identified four primary behavioral traits which he called Dominance (now called Drive), Inducement (now Influence), Submission (now Support, or sometimes Steadiness), and Compliance (now Clarity or Conscientiousness) . The initials D-I-S-C give the model its name. 

To arrive at these traits, Marston theorized that human behavior could be understood along two dimensions: 

  1. How actively a person engages with the world around them. We call this the Active vs. Receptive axis. An “Active” person is decisive and takes action quickly, while a “Receptive” person is more considered in their approach. Drive and Influence types represent the Active side of the spectrum, and Support and Clarity types represent the Receptive side. 
  2. Whether they expect that world to be more challenging or more cooperative. We call this the Skeptical vs. Agreeable axis. Drive and Clarity types represent the Skeptical side of the spectrum; their default posture is task-focused and results-oriented, treating the environment as something to be managed or overcome. Influence and Support types represent the Agreeable side. These types are people-focused and trusting, and they treat the environment as broadly cooperative. 

Plot those two dimensions against each other and Marston’s four behavioral styles emerge

  • D, Drive: An assertive, results-driven style. Drive types approach challenges head-on and gravitate toward decisive, take-charge leadership. 
  • I, Influence. An enthusiastic, people-oriented style. Influence types are excellent team players, natural networkers and great at boosting morale. 
  • S, Support. A patient, cooperative style. Support types tend to be steady, reliable and supportive contributors who value harmony and consistency. 
  • C, Clarity. A precise, detail-oriented style. Clarity types bring careful, thoughtful analysis to their work and prefer accuracy over speed. 

The Statistical Foundation of a Good DISC Assessment 

It may sound obvious, but the starting point for any personality assessment is to know what it is actually assessing. These traits are called constructs and, since a DISC assessment is based on an established theory of personality, the constructs are defined for us — the test must measure where an individual sits on the Active vs. Receptive and Skeptical vs. Agreeable axes, leading us towards a typing of D, I, S or C. 

Truity goes a step further and also measures two additional constructs to help differentiate nuances between opposite types. Specifically, we look at the Drive vs. Support construct, contrasting the dominant, assertive character of the Drive type with the gentle, responsive character of the Support type, and the  Influence vs. Clarity construct, contrasting the relational, enthusiastic nature of the Influence type with the detail-oriented, reserved nature of the Clarity type. The results of this analysis are not mentioned in the test-taker’s type report, but we use them behind the scenes to ensure greater accuracy in scoring.

A DISC assessment might run to 30-40 questions, which may seem excessive. However, a statistically credible assessment will always ask several questions about the construct it is trying to measure, before combining the answers into a single score. Repeating questions in different ways helps see if the test-taker is answering honestly or just guessing or not paying attention. 

With an assessment clear on what it is measuring, it then needs to establish reliability and validity.

Reliability: Does the Test Produce Consistent Results?

The first technical hurdle any assessment has to clear is reliability, meaning that the test measures its constructs consistently. Psychometricians use a statistic called Cronbach's alpha to check how well a test’s items correlate with one another, or whether some of them were accidentally measuring something else and muddying the result. It’s similar to how you might test someone's fitness with five different exercises: push-ups, a sprint, a plank hold, a jump, and a stretch test. If someone who is fit tends to do well on all five, and someone who's unfit tends to do poorly on all five, the five exercises are all capturing the same underlying thing: fitness. But if the stretch test has little to do with how someone does on the other four, it isn't really measuring fitness, and it drags down the accuracy of the combined score. 

Cronbach's alpha is the number that tells you how well a set of questions agree with each other in this way. Alpha scores range from zero to one, with higher scores suggesting greater internal consistency. A result above 0.70 is generally considered a solid benchmark for a DISC assessment, meaning the questions behind each trait are consistently pulling together.

Another form of reliability is “test-retest.” This involves having the same person take the test at two different times to see if scores stay stable, such that if a team member takes a DISC assessment in January and then again in October, perhaps even after they’ve moved to a new team, role or employer, we would expect the results to be roughly the same.

For a manager relying on the results of a DISC assessment to make decisions about coaching or team composition, reliability is a baseline requirement. A test with weak reliability might tell a different story depending on someone's mood that day, which makes it worthless as a basis for any lasting decision.

Validity: Do the Results Actually Predict Anything?

Reliability confirms a test is consistent, but consistency on its own doesn't tell you much. A stopped watch is perfectly consistent, it shows the same time every time you look at it, but that doesn't make it useful. Validity is the check that matters to a manager: does the assessment measure what it claims to measure, and do the results correspond to the observed ways that people work?

This is where a lot of DISC tools fall short, because validity requires collecting outcome data and checking whether assessment scores correlate with it. That takes a large sample and careful analysis, typically using machine learning, using various statistical methods:

  • Content validity looks at whether the individual questions written for each construct actually hang together, like the fitness example we gave above. Researchers compare how test-takers answer the different questions intended to measure the same trait, checking that those answers correlate closely with each other, while correlating much less with questions from a different trait. This confirms the questions are doing the job they were designed for, rather than accidentally blurring one trait into another.
  • Factorial validity takes a step back from the test’s individual questions and asks a more fundamental question: does the underlying data support the existence of four distinct constructs at all, or did the test's designers simply assume that structure in advance? This uses a statistical technique called factor analysis, which looks at patterns across thousands of responses and checks whether they naturally group into four coherent clusters that match the intended DISC constructs, rather than three, or five, or some other number entirely.
  • Predictive validity is the most practically important test for a workplace tool, because it checks whether assessment scores correspond to real behavior outside the test itself. This typically means comparing DISC results against outcomes like a person's actual role, their management responsibilities, or their reported working style, to see whether the assessment predicts actual behavior on the ground rather than just producing a plausible-sounding label. For example, Drive types typically report managing significantly larger teams on average than any other type, along with the highest self-rated level of decision-making responsibility at work. 

What to Ask Before You Trust a DISC Tool

Given all of this, a manager evaluating a DISC assessment for their organization has a few concrete questions to ask their vendor: 

  • Has the publisher produced technical documentation, describing its reliability statistics for each scale, and do those numbers clear the standard threshold? 
  • Has the tool been validated against real-world outcomes with a sample large enough to trust, or does it rely on face validity alone, meaning the results simply sound plausible? 
  • Is the underlying model built from genuine psychological constructs, or does it sort people into categories without a clear theoretical foundation?

A tool that can answer all three affirmatively, and can show its work, is one worth building team structures and coaching programs around. A tool that cannot is, at best, an icebreaker for a team offsite.

Ready to see what a gold-standard DISC assessment looks like? 

Trusted by Fortune 500 companies, Truity’s industry-leading DISC assessment is built on a rigorous technical foundation. Using a modern, diverse sample of over 43,500 test-takers, statistically sound constructs, and validation research tying results to observable outcomes like team size and decision-making responsibility, the assessment has been shown to have excellent reliability and real-world correlations with key workplace outcomes. 

We also recognize that people have a primary DISC type plus a strong secondary type that influences their main DISC style in some way. Our DISC assessment offers a more detailed classification of 12 work styles in total, the four primary types plus eight combination blends (e.g. D with I, I with S and so on) to provide a more nuanced profile of how team members work, communicate and make decisions. 

Truity@Work is our one-stop platform for testing your entire team. To learn more, purchase tests, and set your team up for sustainable success, visit Truity@Work’s DISC Assessment for Teams

Jayne Thompson
Jayne is a B2B tech copywriter and the editorial director here at Truity. When she’s not writing to a deadline, she’s geeking out about personality psychology and conspiracy theories. Jayne is a true ambivert, barely an INTJ, and an Enneagram One. She lives with her husband and daughters in the UK. Find Jayne at White Rose Copywriting.