Posts by Category

agile-ai

Is Agile Failing in the Age of AI?

6 minute read

Published:

Agile transformed software development by adapting to changing requirements. But what happens when AI dramatically reduces the cost of implementation? Are we still optimizing for the right bottleneck?

application-generation

cpu-vs-gpu-battle

CPU vs GPU battle, Round 1: leave the laptop, the ceiling comes with you

18 minute read

Published:

Three rounds of this experiment lived on a 16 GB laptop and nothing broke 0.5 QWK - a lexical floor at 0.222, and every learned or prompted arm crowded between 0.24 and 0.49. So we left the machine: the same prompt, the same parser, the same 3,000 held-out pairs, handed to a 7B judge on a datacenter GPU. It scored ESCI QWK 0.361 and WANDS 0.354 - better than the 3B it scaled up from, and still inside the same band. The ceiling is the task, not the model size. The result is now a poster at SCD 2026.

CPU vs GPU battle, Round 1: adapt the generalist, read the logits

16 minute read

Published:

Part 2 trained a specialist cross-encoder that won in-domain but leaned on its training data. This round we take the generative Llama-3.2-3B, LoRA-tune it on the laptop GPU with MLX, and grade by reading the label-token logits instead of prompting for JSON. In-domain it lands about even with the specialist (ESCI QWK 0.353), but it is the only arm that does not degrade out of domain - WANDS QWK 0.486, up from its own ESCI, while every other arm falls.

CPU vs GPU battle, Round 1: train the judge, don’t prompt it

8 minute read

Published:

In Part 1 we prompted a general model to grade e-commerce search relevance and watched it miscalibrate. This round we stop prompting and train a specialist - a bge-reranker cross-encoder fine-tuned on Amazon ESCI, on the laptop’s own Apple-Metal GPU. It becomes the first arm to beat the shipped on-device judge on both test sets (ESCI QWK 0.360, WANDS 0.299) at ~28 ms/pair in ~1 GB of RAM.

CPU vs GPU battle, Round 1: the agents brought the GPU

13 minute read

Published:

How well does a small prompted LLM grade e-commerce search relevance on a 0-3 scale, entirely on a laptop? We start the CPU-vs-GPU battle with the two easy arms - BM25 and a prompted local model - and find the honest surprise is not accuracy but calibration and circularity.

enterprise-ai

modern-software-small-businesses-nonprofits

MS-SBN, Part 3: My Ten-Year-Old Built a Video Game From Four-Word Prompts

13 minute read

Published:

I typed two sentences into a coding agent, asked it to write a plan, and handed the laptop to my ten-year-old son. Over four weeks Mathai built a real, playable 2D game without writing a line of code. It was a deliberate experiment: proof that when every step has to verify it works, the software keeps working — whoever is driving. That is the core of Mission-Driven Engineering.

MS-SBN, Part 2: Amy Turned a High-School Dream Into a Website

7 minute read

Published:

Amy Givens owned the MyAmy Designs domain since high school, but the website stayed a dream for years. With ChatGPT and Codex, she finally turned that domain into a real business website for MyAmy Designs & Remodeling.

SLAB-RS, Part 8: The Word the Embeddings Couldn’t Read

9 minute read

Published:

Phase 3 measured the whole classical toolkit and shipped almost none of it. One thing did ship — and it didn’t come from a bigger model. It came from reading a thousand failures and noticing that a million-vector embedding model cannot read the word ‘without’. The fix was a single linguistic rule. +4.8 nDCG on negation queries, live.

SLAB-RS, Part 7: The Textbook Wins the Benchmark, and Loses the Trade

10 minute read

Published:

Retail is the first dataset in the project with real graded labels — so it is the first place the classical retail-search toolkit gets a fair audition. Field weighting stayed dormant. Learning-to-rank finally woke up, and then relearned the hybrid it was supposed to beat. The full textbook cascade won the benchmark — by +0.13% over the lean stack that was already shipped. Measure everything, ship almost nothing.

SLAB-RS, Part 6: The Agent Meets a Million Real Products

11 minute read

Published:

Five parts of a retail search series, and the documents were never products. Phase 3 changes that: 1.2 million real Amazon products, graded Exact/Substitute/Complement/Irrelevant labels, and one question — does the academic stack survive contact with real retail? It does. The BGE hybrid lands at parity with the published single-model baseline, at production latency, and int8 quantization makes a live million-vector index fit on a free-tier box, losslessly.

SLAB-RS, Part 5: How the Agents Discovered Hybrid Search

14 minute read

Published:

The agents took their Phase 1 wins to fifteen new domains — and discovered hybrid search: the one method that improved every single one, by up to 64%. The keyword tricks that looked brilliant on aeronautics turned out to be aeronautics-shaped, and even the cleverest ranker got demoted by Occam’s razor.

SLAB-RS, Interlude: Why the Agent Is Not in a Store Yet

16 minute read

Published:

Four articles into a series called retail search and there is not a single product — only aeronautics abstracts. Here is the map: what BEIR is, why the same BM25 scores 0.158 on one dataset and 0.789 on another, why I was wrong to call Phase 2 a filter, and what has to happen before the agent reaches real products.

SLAB-RS, Part 4: The Agent Breaks Its Ceiling with Embeddings

10 minute read

Published:

Embeddings from a chat model scored five times worse than random vectors. A purpose-built retrieval model beat the keyword ceiling by 8.4%. The agent’s three embedding attempts, a machine-learned ranking detour, and a live OpenSearch vector index.

SLAB-RS, Part 3: The Agent Climbs Keyword Search to Its Ceiling

11 minute read

Published:

An AI agent ran five ranking experiments on a live OpenSearch baseline in six days — work that would have taken a core search team months. This article shows the process, the failures, and how a rarely-shipped old technique won the keyword round.

software-scarcity

Software Is Getting Cheaper. But By How Much?

8 minute read

Published:

I’ve spent months arguing AI is making software cheaper. So I went looking for the evidence, applied my own rule, and found something more interesting than a percentage: software purchasing power.

The End of Software Scarcity, Part 2: Rosa’s Story

3 minute read

Published:

In my previous article, I argued that AI-assisted development may be bringing an end to software scarcity. This article tests that idea through a real-world project with a local small business owner.

teaching

Teaching Philosophy as a Part-Time Teacher

3 minute read

Published:

Between 2017 and 2025, I consistently taught undergraduate or graduate courses in most semesters while balancing my teaching responsibilities with a full-time position in industry.

Teaching philosophy

1 minute read

Published:

Teaching methods should be flexible, evolving based on the students, course content, and learning environment.