Has Everybody Forgotten About GIGO?

Most programmers are familiar with the concept of garbage in, garbage out.  The term has been around since the early days of computing with some even attributing it to Charles Babbage and his 1820’s mechanical “difference engine”. In simple terms, if an algorithm takes flawed data as input then the results it produces will also be flawed, even if the algorithm itself is correct. 

‘Flawed data’ is of course not an absolute term. Most datasets, especially those not contrived in a laboratory or cherry picked for academic research, come with their quirks and errors. This error can arise from sampling issues, measurement problems, or old-fashioned transcription errors.

Not too many years ago, computing resources were both minimal and expensive. Algorithms were tested on small batches of data and analyzed thoroughly to make sure the results were as expected. If a dataset was small, it was manually checked at least twice before being used. Statistics programs always kicked out a massive volume of diagnostics specifically so that the analyst could review them looking for anomalies, bad data points, and other signs of trouble. We all knew that one bad data point could render a regression model worthless.

How things have changed. Computing speeds have exploded, and memory and storage have become dirt cheap. In the last couple of years, the improvements in large language models have made it possible to imagine that many of the complex spatial decisions that companies make on a daily basis can be done less expensively and in fact better than the human analysts of today.

We have recently seen at least a dozen startup companies advertising that they have “the solution” to all your location problems. Most talk extensively about the wizardry of their AI systems and say little about the data that is used to feed them. Often, they are staffed by personnel with computer science and applied mathematics backgrounds who have little or no experience in the realm of geospatial data and analysis. If data is mentioned, it is only in passing. These systems might use foot traffic, traffic counts, demographics and a host of other data elements, but they rarely say where they are getting or why they chose them.

Data, it seems, is no longer relevant.  After all, the LLM will sort it all out. 

Or will it?

AI environments are really good at extracting and summarizing simple patterns from data, no question about that. Spatial analysis is at best primitive, although this will no doubt improve dramatically over time. That said, these systems are nowhere near the point where they can overcome deficiencies in the data itself.  

With demographic data, the tendency is to use “census data”, mainly because it doesn’t add to their per user cost structure. How many of these companies are aware that the census data they are using based on a 1% sample of households over a rolling five year period? We are not saying that the ACS data is garbage in, garbage out. It is a great data source which we use extensively.

But here is the problem. The ACS is at best two years out of date and based on data which was collected between 2 and 7 years ago. For most locations, you won’t find a substantial difference between the ACS data and the AGS data, even though we add numerous detailed additional sources. How can you tell if the site you are looking at is one which the ACS substantially underestimates?  Or worse, is declining and the ACS has not captured that decline. Take 1000 block groups. In most cases, the differences will be relatively small and not likely to change a business decision.  But the five percent which do have major differences are dispersed in space and will affect many trade areas. So how do you know if your particular area of interest includes some of the block groups which are drastically different? Or, that another area doesn’t look promising because it has missed all the growth in the last two years?

The AI will, with complete confidence, report its findings while being blissfully unaware of the issue. Those startup companies that tried to shave a few hundred dollars off the annual subscription price will quickly discover that using anything but the best available data is a recipe for failure. When the AI confidently advises opening a store at a bad location, and these companies have no good answer as to why, customers will quickly disappear and go back to the methods which have worked for decades.

Yes, AI is powerful, but no amount of power will overcome inferior data. Garbage in, garbage out is just as relevant in the AI era as it was in the IBM punch-card era.

more blogs

September Round-up

At the end of each month, the AGS team looks back at articles and blog posts that we saw this month that stood out to

Read more >