Close Menu
    Facebook LinkedIn YouTube WhatsApp X (Twitter) Pinterest
    Trending
    • These Were My Favorite Things Samsung Unpacked During Its 2026 Galaxy Event
    • AI minister role boosted but tech department axed in Burnham shake-up
    • Loop Engineering for RAG Question Parsing: The Small Loop That Runs Before Retrieval
    • The risk of weather data sabotage is rising
    • Hand-E Now Reaches 100 mm Without Giving Up an Ounce of Precision
    • Weight loss drug effectiveness and long term maintenance
    • Here’s what Albo’s ‘Office of AI’ means for Australian tech
    • YouTube and X Have Become ‘Gateways’ to Nudify Apps
    Facebook LinkedIn WhatsApp
    Times FeaturedTimes Featured
    Thursday, July 23
    • Home
    • Founders
    • Startups
    • Technology
    • Profiles
    • Entrepreneurs
    • Leaders
    • Students
    • VC Funds
    • More
      • AI
      • Robotics
      • Industries
      • Global
    Times FeaturedTimes Featured
    Home»Artificial Intelligence»Why Your ML Model Works in Training But Fails in Production
    Artificial Intelligence

    Why Your ML Model Works in Training But Fails in Production

    Editor Times FeaturedBy Editor Times FeaturedJanuary 14, 2026No Comments8 Mins Read
    Facebook Twitter Pinterest Telegram LinkedIn Tumblr WhatsApp Email
    Share
    Facebook Twitter LinkedIn Pinterest Telegram Email WhatsApp Copy Link


    , I labored on real-time fraud detection programs and advice fashions for product firms that seemed wonderful throughout improvement. Offline metrics have been sturdy. AUC curves have been secure throughout validation home windows. Characteristic significance plots advised a clear, intuitive story. We shipped with confidence.

    A couple of weeks later, our metrics began to float.

    Click on-through charges on suggestions started to slip. Fraud fashions behaved inconsistently throughout peak hours. Some selections felt overly assured, others oddly blind. The fashions themselves had not degraded. There have been no sudden knowledge outages or damaged pipelines. What failed was our understanding of how the system behaved as soon as it met time, latency, and delayed reality in the actual world.

    This text is about these failures. The quiet, unglamorous issues that present up solely when machine studying programs collide with actuality. Not optimizer decisions or the most recent structure. The issues that don’t seem in notebooks, however floor at 3 a.m. dashboards. 

    My message is straightforward: most manufacturing ML failures are knowledge and time issues, not modeling issues. If you don’t design explicitly for the way data arrives, matures, and modifications, the system will quietly make these assumptions for you.

    Time Journey: An Assumption Leak

    Time journey is the most typical manufacturing ML failure I’ve seen, and likewise the least mentioned in concrete phrases. Everybody nods once you point out leakage. Only a few groups can level to the precise row the place it occurred.

    Let me make it express.

    Think about a fraud dataset with two tables:

    1. transactions: when the fee occurred
    The transactions desk reveals a consumer making a number of funds on December twenty fourth, all earlier than mid afternoon.(Picture by creator, generated utilizing artificial knowledge for illustration)
    1. chargebacks: when the fraud final result was reported
    The chargeback desk reveals a fraud report arriving at 6:40 PM the identical day.(Picture by creator, generated utilizing artificial knowledge for illustration)

    The characteristic we would like is user_chargeback_count_last_30_days.

    The batch job runs on the finish of the day, simply earlier than midnight, and computes chargeback counts for the final 30 days. For consumer U123, the rely is 1. As of midnight, that’s factually appropriate.

    Picture by creator, generated utilizing artificial knowledge for illustration

    Now have a look at the ultimate joined coaching dataset.

    Morning transactions at 9:10 AM and 11:45 AM already carry a chargeback rely of 1. On the time these funds have been made, the chargeback had not but been reported. However the coaching knowledge doesn’t know that. Time has been flattened.

    That is the place the mannequin cheats.

    Picture by creator, generated utilizing artificial knowledge for illustration

    From the mannequin’s perspective, dangerous wanting transactions already include confirmed fraud indicators. Offline recall improves dramatically. Nothing appears to be like mistaken at this level.

    However in manufacturing, the mannequin is meant to by no means sees the long run.

    When deployed, these early transactions should not have a chargeback rely but. The sign disappears and efficiency collapses. 

    This isn’t a modeling mistake. It’s an assumption leak.

    The hidden assumption is {that a} every day batch characteristic is legitimate for all occasions on that day. It’s not. A characteristic is barely legitimate if it may have existed on the actual second the prediction was made. 

    Each characteristic should reply one query:

    “May this worth have existed on the actual second the prediction was made?”

    If the reply isn’t a assured sure, the characteristic is invalid. 

    Characteristic Defaults That Develop into Indicators

    After time journey, it is a quite common failure purpose that I’ve seen in manufacturing programs. Not like leakage, this one doesn’t depend on the long run. It depends on silence.

    Most engineers deal with lacking values as a hygiene drawback. Fill them with common, median or another imputation method after which transfer on. 

    These defaults really feel innocent. One thing secure sufficient so the mannequin can preserve working.

    That assumption seems to be costly.

    In actual programs, lacking hardly ever means random. Lacking usually means new, unknown, not but noticed, or not but trusted. Once we collapse all of that right into a single default worth, the mannequin doesn’t see a niche. It sees a sample.

    Let me make this concrete.

    I first bumped into this in an actual time fraud system the place we used a characteristic referred to as avg_transaction_amount_last_7_days. For energetic customers, this worth was properly behaved. For brand spanking new or inactive customers, the characteristic pipeline returned a default worth of zero.

    Picture by creator, generated utilizing artificial knowledge for illustration

    For instance how the default worth grew to become a robust proxy for consumer standing, I computed the noticed fraud price grouped by the characteristic’s worth:

    knowledge.groupby("avg_txn_amount_last_7_days")["is_fraud"].imply()

    As proven, customers with a price of zero exhibit a markedly decrease fraud price—not as a result of zero spending is inherently secure, however as a result of it implicitly encodes “new or inactive consumer.”

    All customers with a median transaction quantity of zero are non fraud. Not as a result of zero is inherently secure, however as a result of these customers are new/inactive. The mannequin doesn’t study “low spending is secure”. It learns “lacking historical past means secure”.

    The default has turn out to be a sign.

    Throughout coaching, this appears to be like good as precision improves. Then manufacturing visitors modifications.

    A downstream service begins timing out throughout peak hours. Out of the blue, energetic customers quickly lose their historical past options. Their avg_transaction_amount_last_7_days flips to zero. The mannequin confidently marks them as low danger.

    Skilled groups deal with this in another way. They separate absence from worth, monitor characteristic availability explicitly. Most significantly, they by no means permit silence to masquerade as data.

    Inhabitants Shift With out Distribution Shift

    This failure mode took me for much longer to acknowledge, principally as a result of all the standard alarms stayed silent.

    When individuals speak about knowledge drift, they normally imply distribution shift. Characteristic histograms transfer. Percentiles change. KS tests gentle up dashboards. Everybody understands what to do subsequent. Examine upstream knowledge, retrain, recalibrate.

    Inhabitants shift with out distribution shift is completely different. Right here, the characteristic distributions stay secure. Abstract statistics barely transfer. Monitoring dashboards look reassuring. And but, mannequin conduct degrades steadily.

    I first encountered this in a big scale funds danger system that operated throughout a number of consumer segments. The mannequin consumed transaction degree options like quantity, time of day, machine indicators, velocity counters, and service provider class codes. All of those options have been closely monitored. Their distributions barely modified month over month.

    Nonetheless, fraud charges began creeping up in a really particular slice of visitors. What modified was not the information. It was who the information represented.

    Over time, the product expanded into new consumer cohorts. New geographies with completely different fee habits. New service provider classes with unfamiliar transaction patterns. Promotional campaigns that introduced in customers who behaved in another way however nonetheless fell inside the similar numeric ranges. From a distribution perspective, nothing seemed uncommon. However the underlying inhabitants had shifted.

    The mannequin had been educated totally on mature customers with lengthy behavioral histories. Because the consumer base grew, a bigger fraction of visitors got here from newer customers whose conduct seemed statistically comparable however semantically completely different. A transaction quantity of two,000 meant one thing very completely different for an extended tenured consumer than for somebody on their first day. The mannequin didn’t know that, as a result of we had not taught it to care.

    Inhabitants shift with out distribution shift

    See this determine above. It reveals why this failure mode is troublesome to detect in follow. The primary two plots present transaction quantity and short-term velocity distributions for mature and new customers. From a monitoring perspective, these options seem secure with the overlap. If this have been the one sign accessible, most groups would conclude that the information pipeline and mannequin inputs stay wholesome.

    The third plot reveals the actual drawback. Although the characteristic distributions are practically similar, the fraud price differs considerably throughout populations. The mannequin applies the identical choice boundaries to each teams as a result of the inputs look acquainted, however the underlying danger isn’t the identical. What has modified isn’t the information itself, however who the information represents.

    As visitors composition modifications by means of development or enlargement these assumptions cease holding, though the information continues to look statistically regular. With out explicitly modeling inhabitants context or evaluating efficiency throughout cohorts, these failures stay invisible till enterprise metrics start to degrade.

    Earlier than You Go

    Not one of the failures on this article have been brought on by unhealthy fashions.

    The architectures have been affordable. The options have been thoughtfully designed. What failed was the system across the mannequin, particularly the assumptions we made about time, absence, and who the information represented.

    Time isn’t a static index. Labels arrive late. Options mature inconsistently. Batch boundaries hardly ever align with choice moments. Once we ignore that, fashions study from data they are going to by no means see once more.

    If there’s one takeaway, it’s this: sturdy offline metrics will not be proof of correctness. They’re proof that the mannequin matches the assumptions you gave it. The true work of machine studying begins when these assumptions meet actuality.

    Design for that second.

    References & Additional Studying

    [1] ROC Curves and AUC (Google Machine Studying Crash Course)
    https://developers.google.com/machine-learning/crash-course/classification/roc-and-auc

    [2] Kolmogorov–Smirnov Check (Wikipedia)
    https://en.wikipedia.org/wiki/Kolmogorov%E2%80%93Smirnov_test[3] Information Distribution Shifts and Monitoring (Huyen Chip)
    https://huyenchip.com/2022/02/07/data-distribution-shifts-and-monitoring.html



    Source link

    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    Editor Times Featured
    • Website

    Related Posts

    Loop Engineering for RAG Question Parsing: The Small Loop That Runs Before Retrieval

    July 19, 2026

    How to Find the Optimal Coding Agent Interface

    July 9, 2026

    I Completed Five Years in Analytics Consulting: 5 Lessons That Changed How I Work

    June 29, 2026

    GPU-Resident Top-K for Agentic RAG: I Built a CUDA Kernel So My Retrieval Step Would Stop Bouncing Off the GPU

    June 19, 2026

    Can Machine Learning Predict the World Cup?

    June 9, 2026

    Automate Writing Your LLM Prompts

    June 5, 2026

    Comments are closed.

    Editors Picks

    These Were My Favorite Things Samsung Unpacked During Its 2026 Galaxy Event

    July 22, 2026

    AI minister role boosted but tech department axed in Burnham shake-up

    July 21, 2026

    Loop Engineering for RAG Question Parsing: The Small Loop That Runs Before Retrieval

    July 19, 2026

    The risk of weather data sabotage is rising

    July 18, 2026
    Categories
    • Founders
    • Startups
    • Technology
    • Profiles
    • Entrepreneurs
    • Leaders
    • Students
    • VC Funds
    About Us
    About Us

    Welcome to Times Featured, an AI-driven entrepreneurship growth engine that is transforming the future of work, bridging the digital divide and encouraging younger community inclusion in the 4th Industrial Revolution, and nurturing new market leaders.

    Empowering the growth of profiles, leaders, entrepreneurs businesses, and startups on international landscape.

    Asia-Middle East-Europe-North America-Australia-Africa

    Facebook LinkedIn WhatsApp
    Featured Picks

    My Virtual Avatar No Longer Looks Terrible in the Apple Vision Pro

    June 12, 2025

    Our Favorite Motorola Smartphone Is $100 Off

    October 9, 2025

    Galaxy S25 Ultra Review: Greatest Phone Screen Ever, but Let’s Not Talk About the AI

    February 5, 2025
    Categories
    • Founders
    • Startups
    • Technology
    • Profiles
    • Entrepreneurs
    • Leaders
    • Students
    • VC Funds
    Copyright © 2024 Timesfeatured.com IP Limited. All Rights.
    • Privacy Policy
    • Disclaimer
    • Terms and Conditions
    • About us
    • Contact us

    Type above and press Enter to search. Press Esc to cancel.