IQ Tests: What They Measure and What They Miss

IQ Tests: What They Measure and What They Miss

The intelligence quotient has a peculiar cultural status: simultaneously treated as the definitive, scientific measure of raw mental horsepower and dismissed, often by the same people in different contexts, as a crude, culturally biased number that misses everything important about real intelligence. Both reactions contain a piece of the truth. IQ tests measure something real, stable, and practically consequential — and they were never designed to, and do not, capture the full range of traits that make up what most people mean when they talk about genius.

What an IQ Test Actually Contains

Modern IQ tests, such as the widely used Wechsler Adult Intelligence Scale (WAIS) and Wechsler Intelligence Scale for Children (WISC), are not single tests but batteries of multiple subtests, each measuring a somewhat distinct cognitive ability — verbal comprehension, perceptual reasoning (often measured through pattern-completion and spatial tasks), working memory, and processing speed, among others in different test versions. These subtest scores are combined into a composite “Full Scale IQ” score, standardized so that a score of 100 represents the average performance of the general population at a given age, with a standard deviation typically set at 15 points — meaning roughly two-thirds of people score between 85 and 115, and scores above 130 or below 70 represent statistical outliers occurring in only about 2% of the population at each extreme.

The General Factor: g

The single most important and most extensively validated finding underlying modern IQ testing is the existence of what psychologists call “g,” the general intelligence factor. This finding, first identified by psychologist Charles Spearman in 1904 through a statistical technique he developed called factor analysis, describes a robust and repeatedly replicated empirical observation: performance across many different, seemingly unrelated cognitive tasks — vocabulary, spatial reasoning, arithmetic, pattern recognition — tends to correlate positively with each other. People who perform well on one type of cognitive task tend, on average, to perform well on others, even tasks that seem to draw on quite different specific skills. This positive correlation across diverse tasks is what g statistically captures, and it remains, more than a century after Spearman’s original work, one of the most robustly replicated findings in all of psychology.

Importantly, g is a statistical construct derived from patterns of correlation across test performance — it is not, in itself, a direct description of any single brain structure or process, though the neuroscience research on frontoparietal connectivity and processing speed discussed elsewhere in this series represents an active and ongoing effort to identify plausible biological correlates of this statistically well-established but biologically underspecified construct.

What IQ Predicts, and How Well

IQ scores are among the most extensively validated predictors in all of psychology when it comes to certain specific, well-studied outcomes. Meta-analyses find that IQ correlates with academic performance at around r = 0.5 to 0.6, with job performance across a wide range of occupations at around r = 0.5 (rising to higher correlations in more cognitively complex jobs and somewhat lower correlations in simpler, more routine jobs), and more modestly with various health and longevity outcomes, an association researchers refer to as “cognitive epidemiology,” with the leading explanation being that higher measured intelligence in youth correlates with better health decision-making and health literacy across the lifespan, alongside a shared association with childhood socioeconomic circumstances.

These are genuinely substantial predictive relationships by the standards of psychological research, and industrial-organizational psychologists have long noted that general cognitive ability tests remain among the single best available predictors of job training success and job performance across occupations, frequently outperforming structured interviews, reference checks, and many other more commonly used hiring tools in rigorous validity studies.

What IQ Does Not Capture

Here is where the popular critique of IQ testing has real and well-supported substance, rather than simply reflecting discomfort with the concept of measurable intelligence differences. IQ tests were explicitly designed, from their earliest incarnation in Alfred Binet’s original 1905 test — created to identify French schoolchildren needing additional educational support, not to rank the general population by inherent worth — to measure a specific, relatively narrow cluster of analytical and academic-style cognitive abilities. They were never designed to, and do not, directly measure creativity, as discussed in this series’ article on divergent thinking; motivation, persistence, or the psychological traits covered in the grit and passion research; emotional intelligence and social skill, covered in a separate article in this series; practical, real-world problem-solving of the kind psychologist Robert Sternberg has extensively studied and termed “practical intelligence,” which his research has found to be only weakly correlated with traditional IQ measures despite being highly predictive of real-world managerial and professional success in his own studies; or domain-specific expertise and knowledge, which, as the working memory and deliberate practice articles in this series discuss at length, plays an enormous role in expert-level performance that a general reasoning test administered without any specific domain training cannot capture.

Dean Simonton’s extensive research on historical genius and eminence has repeatedly found evidence for a “threshold effect”: IQ correlates meaningfully with achievement up to a point, roughly in the range of 120, above which additional IQ points show a substantially weakened, and in some analyses close to negligible, additional relationship with real-world eminence and creative achievement. Beyond this threshold, Simonton’s and others’ research suggests that other factors — including many of the traits explored throughout this series, from openness to experience and divergent thinking style to sustained deliberate practice and simple access to opportunity — become considerably more important than any additional increment of raw IQ in explaining who goes on to achieve genuinely extraordinary, historically significant accomplishment.

Sternberg’s Triarchic Theory and Gardner’s Multiple Intelligences

Two of the most influential theoretical challenges to the standard IQ framework deserve mention here, though each is explored in greater depth elsewhere in this series. Robert Sternberg’s triarchic theory of intelligence proposes that traditional IQ tests primarily capture what he calls “analytical” intelligence, while largely neglecting two other components he argues are equally important to real-world success: “creative” intelligence, the ability to generate novel and useful ideas, and “practical” intelligence, the ability to apply knowledge effectively to real-world, often ill-defined problems, sometimes described as “tacit knowledge” — the kind of situationally-specific know-how that’s rarely explicitly taught but that experienced, successful people in a field consistently demonstrate. Howard Gardner’s theory of multiple intelligences goes further still, proposing that what we call “intelligence” is actually a set of several largely independent capacities — including musical, bodily-kinesthetic, interpersonal, and intrapersonal intelligences, among others — rather than a single unified construct measurable by any one test. Both theories remain genuinely contested within mainstream psychometrics, with critics arguing that Gardner’s framework in particular has not been operationalized into rigorously validated, independently confirmable measurement instruments the way traditional g-based IQ testing has — a distinct and separate article in this series examines this specific debate in more detail.

The Flynn Effect: A Puzzle Embedded in IQ’s Own History

A particularly striking finding that complicates any simple, fixed reading of IQ scores is the “Flynn effect,” named for researcher James Flynn, who documented that raw IQ test scores rose substantially across most tested populations worldwide throughout the twentieth century — by roughly 3 points per decade in many countries, a large enough cumulative shift that a person scoring average by today’s test norms would have scored well above average by the norms used just a few generations earlier. This rise cannot plausibly reflect genuine, generation-over-generation genetic change on any timescale evolutionary biology would support, and instead is generally attributed to a combination of environmental factors — improved nutrition, reduced childhood disease burden, increased educational attainment, and greater general familiarity with the abstract reasoning style that IQ tests specifically demand. The Flynn effect stands as one of the strongest pieces of evidence that IQ scores, whatever stable and heritable trait they partly capture, are also substantially responsive to environmental and generational conditions — a finding that sits in productive tension with, rather than simple contradiction of, the substantial heritability estimates for IQ discussed in this series’ article on the genetics of intelligence.

The Honest Summary

IQ tests measure a real, robustly replicated, and practically consequential general cognitive ability factor with genuine predictive power for academic and occupational outcomes across large populations — this much is not seriously disputed within mainstream psychometric research. But they were designed for, and remain best suited to, a specific and comparatively narrow purpose: predicting analytical and academic-style performance across a general population, not diagnosing or ranking the far rarer, more idiosyncratic, and more domain-specific combination of traits — creativity, drive, opportunity, unusual cognitive style, and sustained deliberate effort — that the rest of this series’ research suggests actually distinguishes genius-level historical achievement from merely above-average intelligence. A high IQ score is neither meaningless nor sufficient. It’s one real, well-measured ingredient in a considerably larger and more complicated recipe.