Your AI has a plumbing problem

Robert De Niro as Harry Tuttle beside exposed ducts in Brazil (1985).
Robert De Niro in Brazil (1985).

Plenty of SMBs and midmarket companies have messy data. That doesn’t mean the business stinks. It just means they’ll have a harder time putting AI to work.

But now that the AI revolution is well underway, there are a few basics to get right. Data cleaning is one of them. Sorting your shit out. Which, depending on how long you’ve been accumulating it, can be a huge fucking pain in the ass.

Think about your garage

You need an old connector cable. You’re pretty sure it’s in there somewhere. You rummage around for fifteen minutes, move a broken printer, open the wrong box twice, and find it. Fine. It was your mess. You remember making it.

Now send a robot in there.

“Get the cable” leaves a few things out. Which cable? For what? The one in the drawer is broken. The one that works is in a box labeled “Christmas,” for reasons that made sense when you put it there.

Maybe five years from now you buy a robot and it handles all of this perfectly. Great. I’ll happily retire the analogy. For now, it helps explain how much context you bring to a job that looks simple from the outside.

Now think about your company. Thousands of SKUs. Years of PDFs, spreadsheets, emails and records scattered across software that doesn’t all talk to each other. People know which version to use, which customer has two accounts, and who to call when the numbers look wrong. How much of that does the software know?

Giving an agent access to the files gets it into the garage. It still needs a way to work out what it’s looking at.

What does “available” mean?

Think about the work behind a self-driving car. A camera image has to become something the system can act on: vehicles, pedestrians, lanes, movement. Waymo’s public datasets give a glimpse of that work, with millions of labeled objects alongside sensor data and maps.

Your inventory system needs its own context. A model may already know what a purchase order looks like. It doesn’t follow that it knows why your warehouse calls something “available” when sales has already promised it to another customer.

You have to settle what the words mean in your business. Then make those meanings available to the software.

Palantir puts a lot of emphasis on its Ontology: a model of the business that connects its data, relationships, logic, actions and permissions. Cleaning up records is part of the preparation. Defining how the business works gives those records meaning.

Start with the nouns: customers, products, orders. Which records refer to the same thing? Add the adjectives: paid, available, overdue. What does each actually mean? Then the verbs: reserve, approve, ship, refund. Who gets to do those things, under what conditions, and when does somebody need to check?

And then there’s all the stuff your team just knows. The exception they make for a particular customer. The supplier who sends quantities in cases instead of units. The approval that happens in a text message. We need to understand those details, decide with you which ones should carry forward, and turn them into rules the software can follow.

A plumber repairing the drain beneath a sink, with a work light and a bucket under the open pipes.
Somebody has to get under the sink.

So what do we actually do?

We can start with an audit and diagnostic. Pick a workflow and show us how it gets done, including the workarounds. We trace the records, find the gaps, and work out what needs fixing before an agent gets involved. You get a map of the workflow, a prioritized repair list and a scope for the first build.

A lot of the work is ordinary software engineering and data science. Connect the systems. Classify the documents. Resolve duplicates. Fix inconsistent units. Make sure people and agents can only access what they’re allowed to. Coding agents help us build and check the software; your team helps us establish what the records should mean.

You’ll hear the term ETL: extract, transform, load. Get the data out of the source system, put it into a form the next system can use, and load it there. Keep track of where it came from. Data mining helps us spot patterns in the mess, such as a supplier whose files keep producing the same error.

AI can help with the cleanup too. It can suggest matches, classify files and flag things that look wrong. When a record is ambiguous, we need a way to get an answer from someone who knows. Then we need to make sure that correction survives the next import.

You don’t have to clean the whole company before anything useful happens. We start with enough of the garage to do the job you care about. If the first job is processing an order, we get the customers, products, stock and approvals for that workflow into shape. Then we test it against the awkward cases your team already knows about.

That’s AI prep. It’s a huge part of what we do.

Mulholland Data Cleaning Corporation. We sort your shit out.

hello@mulholland.ai

Plumber photograph by ocean yamaha, CC BY 2.0. Resized and converted to WebP.

Technical background: IBM on data cleaning, ETL and data mining. Film still: Robert De Niro as Harry Tuttle in Brazil (1985), directed by Terry Gilliam. Cinematography by Roger Pratt. Frame via Film Grab.