One Saturday morning, I got a text message. Our models were giving the same prediction for every article. It didn’t matter what the news said. Every article got the same score. Nothing was wrong with the models. They were exactly the models we had trained.
At that time, we used text classification models. Such models read numbers, and word embeddings turned the words into those numbers. The embeddings lived in Redis, next to our cache. When Redis ran out of space, someone replaced it with a bigger instance and forgot to load the embeddings back. Without them, every article looked the same to the model.
So far, this book has been about one dataset and one model on one laptop. We will never deploy the MPG model, so it can’t show you what happens after training. Let’s take a break from our MPG regression problem for one chapter. (I will use some classification terms. Don’t worry, I will explain them in the classification chapter.)
In a real project, the model depends on data sources, labels, preprocessing code, deployment pipelines, and monitoring, and any of them can break. In this chapter, I walk through that process on the project where I got that text message. I worked at a startup that built supply chain risk management software, and I once described its setup in a conference talk, so the story is public. I was the MLOps engineer. I built the pipelines that deployed the models and kept them running. The machine learning team trained the models.
The process has three parts. Steps 1-5 happen before any model exists, and you move through them in order. Steps 6-9 are a loop you repeat until the model is good enough. Steps 10 and 11 happen after the model ships, and monitoring sends you back to step 1, either to improve the model or to decide it’s time to work on a different problem.
Under all of it run three things that belong to no single step: governance, security, and the infrastructure that makes the loop repeatable.
Before the model
Nothing in the first five steps trains anything. They decide what the model is for, what it learns from, and how you will know whether it works.
1. Problem framing
Supply chain risk management sounds fancy. What does it mean?
Our clients owned factories and transport companies. Their suppliers had facilities all over the world. When something happened near one of those facilities, our clients wanted to know before it hit their own business. A train line closure. A mass traffic accident. A worker protest. A wildfire that evacuates the town where your supplier happens to have a plant.
We tracked news sources for such events. Our risk assessment team read them and decided which clients should get a notification.
There was far too much news for a team to read. So we built models that filtered out the irrelevant articles. Note the direction. The models didn’t pick the important news. They threw away the useless news, and everything else went to people.
Because the models only removed articles, their two kinds of mistakes had very different costs. A relevant article marked as irrelevant never reaches a human, so nobody gets a chance to catch it. The next time anyone hears about that event is when a client calls and says, “This horrible thing happened, and you didn’t warn me!” An irrelevant article that slips through costs an analyst a minute of reading.
The goal was to cut the team’s workload as much as possible without throwing away anything that mattered. Overall accuracy treats both mistakes as equal, so it couldn’t tell us whether we were reaching that goal.
The labeling rule in step 3 and the threshold in step 9 both follow from this difference in costs, so write it down before you touch the data.
2. Data collection
Start with two questions: what data do you have, and what data are you allowed to use?
If your data comes from users, you need their consent, and privacy law decides what you may store and for how long. We used news, so our problem was licensing. Some sources allowed us to classify their articles in production but not to train on them.
That splits your data in two: what the model sees in production and what it may learn from. When the two differ, you score the model on a distribution it never trained on. Know which sources sit on which side. News is repetitive. Many outlets write about the same event, and some just copy each other. Losing a few sources for training didn’t leave us short of data. (The same repetition comes back to bite us in step 5.)
3. Data labeling
A classifier learns to mimic a label. Someone has to provide that label for every article in the training data. Usually, not even for a full article, just for an excerpt.
Automating a manual process puts you in a privileged position. The risk assessment team was already judging every article, every day. All we had to do was store the input and the team’s decision. The label was “relevant to at least one client.” And the rule for disagreements was simple: if at least one person thinks it’s relevant, it’s relevant.
That rule changes the target the model learns. It pushes the labels toward “relevant,” which is what you want when a missed event costs more than a wasted minute. A majority vote would produce a different target, and the model would learn a different filter.
Decide how you resolve disagreements on purpose, and write it down in the labeling guidelines. Then, measure how often your labelers disagree. If two careful people can’t agree on an article, no model will be right about it. For this problem, the disagreement rate is the noise floor from Rule 10.
This only works until the model ships. From then on, the team sees only what the model lets through. Filtered articles never get a label. Every new training set contains only the articles the previous model already approved, and the mistakes you care about most (relevant news thrown away) become invisible in your data.
The fix is cheap. Send a small random sample of the filtered articles to the labelers anyway. It tells you how much relevant news the filter throws away, and it keeps both classes in the training data. Without it, every retraining learns from a slightly more biased sample than the one before.
4. Exploratory data analysis
Everything we did with the cars at the beginning of this book applies here too. With text, we asked different questions than we did with the cars.
We checked how many relevant articles each source produced, because some sources might be useless. We looked for duplicates: is someone reprinting what was already said, only a few hours late? We checked how long a piece of news stays relevant.
A fire can disrupt a supply chain for months, but nearly all the articles about it appear within a few hours. A worker protest is the opposite. The news builds up for weeks before it happens, and the disruption lasts a day or two.
We also counted the data per language and per region. A model or a metric built on a thin slice deserves the same suspicion as the model years with two Japanese cars (Rule 7).
5. Data preparation
Here, you split the data into the training, validation, and test sets. The danger is data leakage, and Rule 9 applies: the split is a modeling decision. Two questions decide it. What counts as the same article? And what question does the model answer?
Deduplicate by event. Reprints and rewrites of one story count as one data point. If you remove only exact copies, the rewritten versions stay and land on both sides of a random split. Your test score then measures how well the model remembers stories it has already read.
Then, split by time. The model will classify tomorrow’s news, so evaluate it on news published after everything it trained on. Leave a gap between the training period and the test period, so a story and its reprint a few hours later can’t land on both sides.
If you must split randomly, assign whole events to one side of the split. Scikit-learn’s GroupShuffleSplit does that when you pass the event ID as groups. Every article of an event lands on the same side.
The loop
Steps 6-9 repeat. You build features, pick a model, tune it, and evaluate it on the validation data. Then you go around again with what the evaluation taught you. When the misses point at the data (wrong labels, a missing source, a leaky split), go back to steps 3-5 and fix the data before you try a bigger model. You leave the loop when the model is good enough for the goal from step 1, or when it’s clear it never will be.
6. Feature engineering
This is where the Saturday text message came from.
Feature engineering builds the inputs the model uses: encodings, aggregations, embeddings, and scaling. Our models used only word embeddings. In production, the backend application turned the text into vectors before calling the model, because TensorFlow Serving runs the model and nothing else. The model lived in one place, the embeddings in another, and the preprocessing code in a third.
The rule: the same logic must run identically in training and in production. The safest way to guarantee it is to ship the preprocessing and the model as one artifact.
In scikit-learn, that’s a Pipeline. When you call fit() on it with the training data, every transformer learns its parameters from that data only. You store the fitted pipeline, and it applies the same steps to every prediction.
In our new setup, I did the same thing at a bigger scale. Each model went into one Docker container, together with its embeddings and its preprocessing code, deployed as a single endpoint. The embeddings shipped with the model, so a running model could no longer be missing them.
We will get to feature engineering for the cars in a later chapter.
7. Model selection
Start simple. Our first models were tiny neural networks: four feedforward layers, fewer than 500 neurons in total. They worked surprisingly well.
A simple model gives you a baseline (Rule 10). Every bigger model has to beat it by a margin large enough to justify the extra cost, because bigger models cost more to train, to serve, and to debug.
We had separate models for each source and language. Over time, we moved to BERT. A single BERT model took more than 1 GB, which pushed us off our old hosting. Then, we started merging the per-language models into one multilingual model. It didn’t take the language as an input. Today, they probably use LLMs. I don’t know. I left around the time ChatGPT went mainstream.
8. Hyperparameter tuning and experiment tracking
Hyperparameters are the settings you choose before training: the number of layers, the learning rate, the regularization strength. You choose them by comparing models on the validation set or on the cross-validation folds. Never on the test set. Every comparison you make on it spends a little of it (Rule 14).
Tuning means dozens or hundreds of runs. A month later, someone asks which data, which code, and which settings produced the model in production. Experiment tracking answers that question. A tool like MLflow records the parameters, the metrics, the code version, and the trained model for every run, automatically.
I deployed MLflow for our ML engineers. They never used it. They tracked their experiments in Excel.
I get the appeal. A spreadsheet needs no setup and bends to whatever you want to write down. For your very first model, it’s fine, because you have nothing to compare it with yet.
But a spreadsheet holds only the numbers someone typed in. When it says one model beat another, someone still has to remember which data and which commit stood behind each row. When that person leaves, the information leaves with them. Once you have a second model to compare against, use the tracking tool.
9. Evaluation
Measure performance on held-out data, with a metric tied to the goal from step 1.
AUC (the area under the ROC curve) is a good choice here. It measures how well the model ranks relevant articles above irrelevant ones across every possible threshold, which makes it a good number for comparing models. With one caveat: it averages over every threshold, including ones you’d never ship. Here, only the high-recall end of the curve matters. And AUC doesn’t decide what happens in production, because production uses a single threshold.
Our thresholds weren’t 0.5. For some languages, the cutoff was 0.85. Each language had its own.
The threshold decides how much relevant news you lose in exchange for how much less reading. Where that balance sits is up to the business, and it can differ between languages. Separate models also produce scores that don’t mean the same thing, so a 0.85 from one model isn’t a 0.85 from another unless you calibrate them. Two models with the same AUC can give very different answers at the threshold you actually ship.
So choose the threshold on validation data, then report what the business cares about at that threshold: the share of relevant articles the filter throws away and the share of the workload it removes. If your model makes decisions about people, add fairness checks across the groups it affects.
The test set stays out of the loop. Look at it once, when you leave the loop, after every decision is made (Rule 14).
After the model
A model that passed evaluation has only proven itself on data you already had. The last two steps get it in front of real traffic and keep watching it once it’s there.
10. Packaging, versioning, and deployment
Store every trained model in a model registry with its lineage: the data, the code, the hyperparameters, and the metrics that produced it. When something goes wrong in production, the registry tells you exactly what is running and lets you roll back to what ran before.
Before a model gets any real traffic, test it. Our ML engineers gave me a set of test cases: specific inputs and the outputs they expected for them. They were unit tests for the model. All of them had to pass, 100%, or the model didn’t ship.
I ran them twice. First, on the trained model in the deployment pipeline, then again on the deployed endpoint, before any traffic went to it. The second run catches what the first one can’t: a broken container, a wrong configuration, missing embeddings.
Those tests check that the code works. Whether the new model is better than the old one shows only on real data, so we released every model in two phases.
First, a shadow deployment. The old model keeps serving every request, the new one gets a copy of the traffic, and nobody uses its answers. Then, a canary release. The new model handles a small share of real requests, and the share grows as long as the results hold. Each phase ran for a few days, while the ML engineers checked whether the new model gave the results they expected on real data.
11. Monitoring
My job was monitoring the infrastructure: latency, memory usage, and traffic. You need that, but it wouldn’t have caught the Saturday problem.
A model that returns the same score for every article answers fast and uses little memory. Every infrastructure metric can look healthy while the model is useless.
So also monitor the model’s behavior. Track the distribution of its inputs (data drift) and the distribution of its scores (prediction drift). A score distribution that collapses to a single value, as ours did that Saturday, is easy to detect.
Neither one tells you whether the model is still right (Rule 13). For that, you need fresh labels, and the random sample of filtered articles from step 3 provides them.
Our ML engineers didn’t wait for a drift signal at all. They worked in a cycle. As soon as they finished their current batch of tasks, they retrained the oldest model. Every model got refreshed in turn, drift or no drift.
There was a reason for that habit. Our oldest models were trained before COVID and knew nothing about it.
Scheduled retraining is a reasonable default with two catches. A retrained model is a new model, so it goes through steps 9 and 10 again, every time. And retraining fixes the drift you didn’t detect without telling you anything about it. You still need drift monitoring to learn when the data changed and by how much.
Under every step
Governance, security, and infrastructure don’t get a step number because every step depends on them. Without them, nobody can say what is running, who changed it, or how to undo it.
Governance and documentation
Every model in production needs a model card: a short document that says what it does, which data trained it, how it was evaluated, and where it is known to fail. Add documentation of the training data and an audit trail of who changed what, and when. They are tedious to write, and you will need them the first time a client asks why an event was filtered out, or when the person who trained the model leaves.
If your model makes decisions about people, regulators may require that documentation. The EU AI Act lists credit scoring and hiring among its high-risk uses and puts obligations on whoever builds and deploys such systems. Check which rules apply before you start building.
Security and privacy
A model is software, and it needs the same access control as any production system: who can read the training data, who can push a model to the registry, and who can change its configuration. Training data often holds more personal data than anyone noticed.
Models also have attacks of their own. If attackers can influence your training data, they can poison it and teach the model a blind spot. If they can query the model freely, they can probe it for the inputs it gets wrong.
MLOps infrastructure
Pipelines, automation, and CI/CD let you repeat the loop without manual work. A deterministic training pipeline gives you the same model from the same data. Automated deployment gives you one button to ship a model and the same button to roll it back.
Don’t build all of it at once. For your first model, you need a place to run it and almost nothing else. Add each piece when the lack of it starts to hurt. That’s how we ended up with a deployment that took under an hour and needed zero human oversight. You clicked a button, and some time later, a message told you whether the model deployed correctly or failed the automated checks. The shadow deployment and the canary release came after that, and judging those was a job for people.