That is submit #2 in my Ivory Tower Notes collection. In submit #1, I wrote about the problem: how each information and AI challenge begins.
This time, the subject is the methodology, and why “immediate in, slop out” is what usually occurs after we skip it.
Immediate in, slop out
I smirked barely when considered one of my connections commented, “You Sent Me AI Slop” underneath some random submit that had tons of of likes. The submit, which contained a call matrix, provided steerage on which platform to make use of for particular information workloads, albeit with questionable standards. High quality apart, it actually seemed nice.
My amusement didn’t finish there as I considered how AIS, i.e., “AI Slop”, must be added as a button to all social media now alongside the like button.
If any YouTube people learn this, it is a function thought as an alternative of quizzing individuals, “Does this feel like AI slop?”
Nonetheless, YouTube nailed the “really feel” half as a result of all of us are likely to make choices primarily based on feelings, usually on the expense of essential pondering.
Why would we make investments power in empiricism, rationalism, and scepticism when now we have AI now? Deadlines usually are not on our aspect, and now we have this new device that delivers outputs for us, whatever the “immediate in, slop out” impact.
However let’s assume you’re genuinely involved in how Platform A compares to Platform B by way of machine studying (ML) capabilities, since you’ve observed two information groups in your organization utilizing separate platforms for nearly an identical ML use circumstances. So, your purpose is to compile an goal overview of each and suggest lowering improvement prices by maintaining just one.
What now? How do you decide whether or not it’s best to consolidate ML workloads?
Certainly not by relying purely on AI, however fairly on…
The trail of inquiry
And so that you’re again to Ivory Tower days once more, the place you have been taught that each discovery is roofed by “The methodology”:
The issue → The speculation → Testing the speculation → The conclusions
Furthermore, you have been taught that discovering the problem is half the work, and the artwork of getting there lies in asking good inquiries to slim it right down to one thing particular and testable.
Therefore, you’re taking the obscure query, “Ought to we consolidate onto one ML platform?”, and you retain rewriting it till it turns into one thing a check can reply:
Does Platform A run our churn pipeline at comparable accuracy and decrease value than Platform B?
Now you may have outlined a topic, a comparability, and issues you may measure, which is sufficient to flip a enterprise query right into a testable speculation.
However first, you do your homework and collect extra info, comparable to what Platform B prices per job at present, what accuracy it hits, and the way it’s designed (e.g., the info, algorithm, and hyperparameters it makes use of), in an effort to reproduce the pipeline on Platform A.
Then, earlier than opinions in your query begin to roll in, you state:
If we run the identical churn pipeline on Platform A as an alternative of Platform B, utilizing an identical information, algorithm and hyperparameters, then the median per-job value will drop by a minimum of 15%, whereas the imply accuracy stays inside 1 share level of Platform B’s.
With this “if-then” formulation, you managed to cool down (a minimum of some) opinionated solutions, realizing that the PoC comes subsequent. Thus, to check the acknowledged presumption, you design and run the PoC, the place you modify solely the impartial variable, which is the platform. Along with this, you freeze the management variables: the dataset, the algorithm, and the hyperparameters, and measure value and accuracy, that are your dependent variables.
You additionally repeat the run a number of occasions to separate the sign from the noise by accumulating a number of information factors, contemplating how a single run will be skewed by environmental noise (e.g., cache), and also you need to keep away from that situation. Then you definitely account for extra nuances, e.g., triggering runs at totally different occasions of day (morning, night, or evening), to show each platforms to the identical mixture of circumstances.
Lastly, you gather all the outcomes and consider the info in opposition to your speculation, which leads you to considered one of these three outcomes:
- Final result 1: The information helps your speculation*. The a number of runs present how Platform A is a minimum of 15% cheaper, and accuracy remained inside the outlined threshold. (*For the notice solely: the info will assist, however not show your speculation, i.e., it provides you with a purpose to carry on to it, which in science is as near a “sure” as you get.)
- Final result 2: The information rejects your speculation. The a number of runs present how Platform A failed to fulfill one or each standards; it was solely 5% cheaper, or the price dropped, however the accuracy degraded past the outlined threshold.
- Final result 3: Your runs are too noisy to name it both means, and the one reply is to maintain testing earlier than drawing any conclusions.
Whichever situation you land in, you may have findings: you both confirmed your educated guess, realized one thing new, or found that that you must maintain testing.
And to be clear about this brief instance: the primary two conclusions received’t provide the inexperienced gentle to consolidate two platforms. Company actuality (and a radical analysis) is a bit messier than that, and there’s extra information (to be collected and evaluated) affecting individuals and processes than a single-scoped PoC can settle.
All proper, we will cease with the methodology now, as a result of most of you’re most likely studying the steps above and questioning…
What the dickens? The place’s AI in all this?
I can solely think about how one thing just like: “MCP, agentic frameworks, brokers,…” was going by means of your head when studying the steps above. Couldn’t agree extra, all great things, and that is how you could possibly pace up the method.
Nonetheless, merely posting AI outputs from a immediate like, “Give me an outline on how Platform A compares to Platform B for ML workloads,” is the place the slop happens, and “if you aren’t doing the hands-on, your opinion about it is very likely to be completely wrong.”
“For those who aren’t doing the hands-on, your opinion about it is vitally prone to be utterly flawed.”
Relevance and constructive affect don’t come from fairly AI posts or presentation infographics, and so they can damage work relationships.
When you find yourself already influencing and need to be seen as an authority, it will be simpler to share views and findings from real-life experiments and your personal confirmed expertise.
As a substitute of beginning your posts “That is the place it’s best to use Platform A over Platform B for…”, attempt one thing extra concrete (if it’s true, in fact):
“Once we (I) modified the [independent variable] to see the way it impacts the [dependent variable], whereas maintaining the [control variables] the identical, our (mine) findings have been…”
After which see whether or not the variety of your followers will increase, and report again the findings.
The inspiration for this submit got here from a Croatian paper by Professor Mladen Šolić, “Uvod u znanstveni rad” (Introduction to Scientific Analysis, 2005, [LINK]). I first learn it as a pupil, and it’s nonetheless one of many clearest explanations of methods to conduct scientific analysis I’ve come throughout.
Thanks for studying.
For those who discovered this submit invaluable, be at liberty to share it along with your community. 👏

