We recently had a potential client ask us an interesting question: Could we provide them with a sample of data from a specific year so they could compare it to the ACS data from the same year? On the surface, the request makes perfect sense. They were essentially asking a very reasonable question: How accurate is your data? Our answer might surprise some people: We don’t necessarily want the data to match.
First, it is important to recognize that no demographic dataset is perfectly accurate, including data from the U.S. Census Bureau. The American Community Survey (ACS) is, by design, an estimate. The Census Bureau surveys a sample of households and uses those responses to estimate characteristics for the larger population. As geographies get smaller and populations are divided into increasingly specific characteristics, margins of error can become substantial.
But sampling error is only part of the issue. For the past six years, we have written extensively about another challenge with Census data: differential privacy. Differential privacy was introduced by the Census Bureau as a way to protect the confidentiality of individual respondents. The basic concept is understandable, and, in an era of increasingly powerful computing and massive amounts of publicly available information, the underlying privacy concern is legitimate. The problem is what happens to the data in the process.
To prevent individuals from potentially being identified, the Census Bureau intentionally introduces statistical noise into certain published data. In other words, some of the numbers are deliberately altered. At large geographic levels, those changes may have relatively little impact. At small geographies—the places where many of our clients actually work—the effects can be much more significant.
We first began raising concerns about differential privacy years ago, as the Census Bureau prepared to implement the methodology for the 2020 Census. Since then, we have continued to examine what it means for demographic data users. The issue is especially important when looking at small areas, small population groups, or very specific combinations of demographic characteristics.
This creates an interesting problem for companies like AGS.
Traditionally, Census data has been treated as the benchmark against which demographic estimates are measured. If your estimate was close to the Census, that was generally considered evidence that your methodology was working.
But what happens when the benchmark itself contains intentional statistical distortion? Simply forcing our estimates to reproduce Census or ACS values would not necessarily make our data more accurate. In some cases, it could do exactly the opposite. Our job is not to replicate a published number. Our job is to produce the best possible estimate of what is actually happening in a community.
That means evaluating a much broader universe of information. Census and ACS data remain important inputs, but they are inputs—not unquestionable ground truth. We can compare them with administrative records, housing and development patterns, consumer and household information, geographic data, historical trends, and other sources that help us understand whether a published estimate makes sense.
If several strong indicators tell us that a community has changed in one direction, we shouldn’t ignore them simply because matching the ACS would make for an easier validation exercise.
This is increasingly important as we look toward the future of demographic data.
For years, we have talked internally about a long-term goal that would have sounded almost radical not that long ago: becoming Census independent. Census independence doesn’t mean ignoring the Census. It means developing estimates and projections that don’t depend on the Census being the definitive answer. In preparation for the 2030 Census, we have been significantly expanding the source data that goes into our models so we can better understand population change independently of what the Census ultimately reports—or what methodology it uses to protect that data.
That work includes incorporating sources such as parcel data, building permits, land-use information, and our identification of “hot blocks” where significant residential growth is occurring. Together, these sources can tell us where housing exists, where new housing is being built, where development is happening, and where population is likely changing—often well before those changes are fully reflected in traditional government estimates.
Will we look at the 2030 Census when it comes out? Of course. Some of the data will still be enormously valuable. But rather than treating it as the answer key that our estimates must conform to, we increasingly see it as another piece of evidence—one part of a much larger puzzle.
So, if you put an AGS estimate next to an ACS estimate from the same year and discover that the numbers don’t perfectly match, that alone doesn’t tell you which one is more accurate. In fact, sometimes the most important question isn’t, “How closely does this match the Census?” It’s “Which number best reflects what is actually happening on the ground?” That’s the question we’re interested in answering—and, increasingly, we believe answering it well means being willing to disagree with the Census.