I’ve spent months arguing AI is making software cheaper. So I went looking for the evidence, applied my own rule, and found something more interesting than a percentage: software purchasing power.
Three rounds of this experiment lived on a 16 GB laptop and nothing broke 0.5 QWK - a lexical floor at 0.222, and every learned or prompted arm crowded between 0.24 and 0.49. So we left the machine: the same prompt, the same parser, the same 3,000 held-out pairs, handed to a 7B judge on a datacenter GPU. It scored ESCI QWK 0.361 and WANDS 0.354 - better than the 3B it scaled up from, and still inside the same band. The ceiling is the task, not the model size. The result is now a poster at SCD 2026.
Part 2 trained a specialist cross-encoder that won in-domain but leaned on its training data. This round we take the generative Llama-3.2-3B, LoRA-tune it on the laptop GPU with MLX, and grade by reading the label-token logits instead of prompting for JSON. In-domain it lands about even with the specialist (ESCI QWK 0.353), but it is the only arm that does not degrade out of domain - WANDS QWK 0.486, up from its own ESCI, while every other arm falls.
In Part 1 we prompted a general model to grade e-commerce search relevance and watched it miscalibrate. This round we stop prompting and train a specialist - a bge-reranker cross-encoder fine-tuned on Amazon ESCI, on the laptop’s own Apple-Metal GPU. It becomes the first arm to beat the shipped on-device judge on both test sets (ESCI QWK 0.360, WANDS 0.299) at ~28 ms/pair in ~1 GB of RAM.
How well does a small prompted LLM grade e-commerce search relevance on a 0-3 scale, entirely on a laptop? We start the CPU-vs-GPU battle with the two easy arms - BM25 and a prompted local model - and find the honest surprise is not accuracy but calibration and circularity.
Phase 3 measured the whole classical toolkit and shipped almost none of it. One thing did ship — and it didn’t come from a bigger model. It came from reading a thousand failures and noticing that a million-vector embedding model cannot read the word ‘without’. The fix was a single linguistic rule. +4.8 nDCG on negation queries, live.
Retail is the first dataset in the project with real graded labels — so it is the first place the classical retail-search toolkit gets a fair audition. Field weighting stayed dormant. Learning-to-rank finally woke up, and then relearned the hybrid it was supposed to beat. The full textbook cascade won the benchmark — by +0.13% over the lean stack that was already shipped. Measure everything, ship almost nothing.
Five parts of a retail search series, and the documents were never products. Phase 3 changes that: 1.2 million real Amazon products, graded Exact/Substitute/Complement/Irrelevant labels, and one question — does the academic stack survive contact with real retail? It does. The BGE hybrid lands at parity with the published single-model baseline, at production latency, and int8 quantization makes a live million-vector index fit on a free-tier box, losslessly.
I typed two sentences into a coding agent, asked it to write a plan, and handed the laptop to my ten-year-old son. Over four weeks Mathai built a real, playable 2D game without writing a line of code. It was a deliberate experiment: proof that when every step has to verify it works, the software keeps working — whoever is driving. That is the core of Mission-Driven Engineering.
The agents took their Phase 1 wins to fifteen new domains — and discovered hybrid search: the one method that improved every single one, by up to 64%. The keyword tricks that looked brilliant on aeronautics turned out to be aeronautics-shaped, and even the cleverest ranker got demoted by Occam’s razor.
Four articles into a series called retail search and there is not a single product — only aeronautics abstracts. Here is the map: what BEIR is, why the same BM25 scores 0.158 on one dataset and 0.789 on another, why I was wrong to call Phase 2 a filter, and what has to happen before the agent reaches real products.
Embeddings from a chat model scored five times worse than random vectors. A purpose-built retrieval model beat the keyword ceiling by 8.4%. The agent’s three embedding attempts, a machine-learned ranking detour, and a live OpenSearch vector index.
An AI agent ran five ranking experiments on a live OpenSearch baseline in six days — work that would have taken a core search team months. This article shows the process, the failures, and how a rarely-shipped old technique won the keyword round.
Before search agents tune anything, the system needs a measurable baseline. This article shows Phase 1 of the project: an out-of-box OpenSearch BM25 baseline on Cranfield, the live relevance numbers, and the failures agents will need to improve.
Amy Givens owned the MyAmy Designs domain since high school, but the website stayed a dream for years. With ChatGPT and Codex, she finally turned that domain into a real business website for MyAmy Designs & Remodeling.
Retail search is not just about returning relevant products. The hard part is deciding which relevant product should come first when customer needs, business goals, inventory, promotions, trust, and mobile behavior all compete.
AI agents and inexpensive hosted services are making the old static-versus-dynamic website distinction less important for small businesses and nonprofits. A website can now become the first useful custom workflow.
If coding agents handle more implementation, the human role does not disappear. It moves toward mission, judgment, validation, integration, accountability, and helping more organizations use software well.
Mission-Driven Engineering did not start as a framework. It started when I realized I was using coding agents to manage UI screens, data layers, and implementation artifacts instead of asking them to satisfy the outcome I actually cared about.
A practical walk-through of Mission-Driven Engineering: how missions, validations, generations, learning loops, and shared MDE memory turn AI coding agents into application-generation systems.
AI application generation feels new, but it echoes Model-Driven Engineering: humans describe intent, machines generate implementation, and independent validation decides whether the result actually works.
AI coding agents can make implementation dramatically faster, but they also create a new bottleneck: the human cost of managing context, attention, and learning across many parallel projects.
A website is only the front door. In this article, I explore how an AI-enabled conversation becomes part of a small operations system for leads, clients, jobs, reviews, videos, and ads.
As AI coding agents improved, the generated code became less interesting than the final application outcome. The question shifted from whether the code looked right to whether the application solved the problem.
Most of my interaction with coding agents became copying error messages from build systems and asking the agent to fix them. That raised an uncomfortable question: why was I in the loop at all?
I initially trusted AI through code generation because code could be validated. What surprised me was discovering that the hardest problem was no longer implementation, but defining when the work was actually complete.
Building the website was easy. The more interesting question was whether a small business could afford an AI-powered customer experience with virtually no recurring software costs.
Agile transformed software development by adapting to changing requirements. But what happens when AI dramatically reduces the cost of implementation? Are we still optimizing for the right bottleneck?
In my previous article, I argued that AI-assisted development may be bringing an end to software scarcity. This article tests that idea through a real-world project with a local small business owner.
For decades, organizations adapted themselves to software because software was expensive to build and maintain. AI-assisted development may reverse that relationship, making it practical for software to adapt to the unique needs of individual organizations.
Moving beyond simple keyword matching. How the industry transitioned to hybrid systems, query understanding, and the “Builder’s Era” of search relevance.
Between 2017 and 2025, I consistently taught undergraduate or graduate courses in most semesters while balancing my teaching responsibilities with a full-time position in industry.