Semantic Models Are Going Away. Really?
- Bill Donofrio

- 22 hours ago
- 7 min read
Updated: 18 hours ago

"Semantic models are going away."
That's what a Chief Product Officer told me a few months ago with complete confidence. He came to this conclusion after a conversation with his AI consultant friend. He said AI is going to make them unnecessary.
He's wrong and the reason he's wrong says more about how little people understand generative AI than it does about semantic models. To the unobservant individual, AI is a magical technology built to do all our thinking for us.
This isn't a semantics argument, pun intended. It's a data argument. AI doesn't run on magic, it runs on probability, and probability needs structure to work with. Kill the structure and you don't get a smarter AI. You get a more confident one that is doing nothing more than guessing.
Let's start with what a semantic model actually is, since that's usually where this conversation goes wrong.
What Is a Semantic Model, Really?
To begin, let's define what a semantic model actually is. Since that definition isn't in the Webster dictionary, I had Gemini provide a concise one, shown below.
A semantic model is a business-facing data layer that translates complex database structures into clear, real-world concepts. It acts as a single source of truth by defining standard business rules, metrics, and relationships, allowing users to query data using familiar terms without needing to know SQL or the underlying table schemas.
Those of us familiar with Power BI may refer to our star schema as the semantic model. However, the definition above is much better. As it states, a semantic model is the "single source of truth." In other words, it's the glue that holds all the data together. It ensures that a sale date means the same thing to the sales team as it does to the finance team. If you sell a house, is the sale considered complete when both sides sign the contract? When the bank approves the loan? Or when the attorney finally hands the new buyer their keys on day one? Eager salespeople like to count the sale before the ink dries on the contract, while accountants cautiously wait until the check clears the bank.
"It's the glue that holds all the data together."
Why a Shared Standard Matters

The reason a semantic model is important is the same reason data governance is still important. It comes down to agreeing on the same standard across the board. Behr Midnight Blue paint is assigned SKU N480-7, with the receipt below for a gallon of paint. Whether you buy that paint in Flagstaff, Arizona, or Chicago, Illinois, the color will be identical when you leave Home Depot.
• Lamp Black: ~5 to 6 oz
• Phthalo Blue: ~3 to 4 oz
• Magenta / Red: ~0.25 to 0.5 oz
• White: ~0.5 to 1 oz
The Imperfect World We Live In

This seems simple to understand at a high level, but it gets far more complicated once you move into subjective results. Take sentiment analysis, for example. A statement like "This car is awesome!" should have a positive score. However, that may not be true if it follows a statement like this one:
"I bought a car this weekend from a used car lot in Tinley Park, and within ten minutes of driving it, I noticed a burning rubber smell coming from one of the belts, a tire reading 20 lbs of pressure, a faint ticking sound from the engine, and a slight pull to the left when braking."
Here's another example. Say a doctor sees a patient who describes their symptoms as widespread pain, fatigue, and joint discomfort. One doctor might diagnose fibromyalgia, another chronic fatigue syndrome, and two more might call it rheumatoid arthritis or lupus. Because the symptoms overlap and there's no X-ray or CT scan that can confirm the diagnosis, the doctor has to work with whatever data is available.
I once had a painful case of shingles that lasted three weeks. The pain sat along a nerve on the back left side of my head, striking every few seconds, and it was excruciating. When I first went to urgent care, the doctor was concerned it might be a neurological disorder. A few days later, a brain scan ruled out any issue with the nervous system but still couldn't identify the root cause. Finally, about a week in, blisters started forming on the affected area of my head. The next doctor I saw asked whether I'd had chicken pox as a child, inspected the blisters, and immediately recognized shingles. He prescribed steroids, and within six days the pain was gone. Just like a doctor needs the right data standard to surface a clear diagnosis, an AI needs a semantic model to prevent a hall of mirrors.
The point is that we don't always have accurate data. Sometimes the data changes over time, revealing new conclusions. Other times it's simply wrong or misleading. This is exactly why data architects, engineers, and analysts are so valuable, and why their jobs can be so stressful.
"We don't always get accurate data. Sometimes it changes over time. Other times it's simply wrong, or misleading."
How AI Actually Works
Since AI is built entirely on probabilities, you want to weed out inaccurate data and feed it accurate data instead. That's why some models perform well and others don't.
"AI doesn't fix bad data. It just makes bad data sound confident."
When you scan the color of a piece of wood, Home Depot can generally match it. Colors are exact, and unless the wood's finish is covered in mud, dust, or grease, the match works. Code generation also performs well in small snippets. Image generation and sentence structuring tend to be far more problematic.
If you give an AI model the definitions below, it should work well:
• Blueprint (S470-5): #7B9BA9
• Midnight Blue (N480-7): #4C5660
• Compass Blue (MQ5-54): #35475B
• Adirondack Blue (N480-5): #748591
• Peaceful Blue (S470-3): #9AB6C0
However, if you provide these descriptions instead, it may not be nearly as helpful.
What exactly is "denim blue"? Jeans come in an entire spectrum of blues.
• Blueprint (S470-5): #7B9BA9, muted denim blue
• Midnight Blue (N480-7): #4C5660, deep slate navy
• Compass Blue (MQ5-54): #35475B, rich regal navy
• Adirondack Blue (N480-5): #748591, soft blue-gray
• Peaceful Blue (S470-3): #9AB6C0, dusty sky gray
Which prompt is more likely to return the right match?
"I'd like a color that's dark blue."
Or:
"Please match to hex #4C5660."
Now consider the accounting software Great Plains. If you're a data engineer who has seen the database behind this product, you're already cringing. Every column name is eight characters long, and table names are cryptic codes made up of a few letters followed by numbers. The only way to make sense of the system is to find a mapping online, or hope someone has been around long enough to decode it.
That link is a list of the commonly used tables out of the hundreds that exist in a real-world implementation. Without that translation, would you have any idea what the GL00102 table actually does? This is exactly why renaming columns and tables to make sense to a human will make your AI tools far more effective.
Since there are limitations on how many words you can feed into a prompt or response, it's also worth considering your chunking strategy. Embeddings built from entire, related sections of information perform better than embeddings created by slicing sentences apart purely by token count. Removing emojis and stray characters helps too. Ask yourself, if you pulled a random chunk out of the database, would it actually make sense to you?
The Walmart Use Case

Let’s use retail giant Walmart as a real-world use case for a strong semantic model. Like most companies they were drowning in a data swamp. Product listings came from third-party sellers and suppliers which were incomplete, inconsistent, and often just wrong. There was no single source of truth. Meanwhile a ton of useful product detail was sitting untouched in titles, descriptions, manuals, and reviews. Prior to the Gen AI revolution, this unstructured data was mostly ignored.
In a nutshell, Walmart needed to take all this new data, structure it, and build appropriate relationships in a graph database that actually understood the context. In their blog post, they explain the frustration of working with the word “cherry”. With thousands of items, was this a scent, flavor, or color of wood stain? The relationship to the actual product category had to be set up correctly. If a chat bot recommended a cherry flavored coffee table, customers would quickly lose confidence. With thousands of product lines and products, Walmart had to be meticulous about data quality. Otherwise, this project would have been doomed from the start.
Final Thoughts
So no, the semantic model isn't dead, and it never will be. Clarity is what lets humans make good decisions, and AI runs on exactly the same requirement. It all boils down to sets and groupings. How you organize your data determines how well everything built on top of it performs in your dashboards, reports, and AI models alike. Your database is like a library full of books. If they're all scattered randomly across the shelves, it becomes nearly impossible to find anything.
Before you blame your AI tool for a bad answer, check three things first. Do your tables and columns actually mean something to a human? Does everyone in the business agree on what your core metrics mean? Is the data feeding the model clean in the first place? Nine times out of ten, that's where the real problem lives.
My friend, Harvard degree and all, was wrong when he declared semantic models dead. You don't have to be. Go look at your own table names right now. If they don't make sense to you, they won't make sense to your AI either.
"If a human can't figure it out, an AI never stood a chance."






Comments