Grounded, production-ready AI — RAG, LLMs and the data behind them.
The hard part of AI in production is not calling a model — it is feeding it the right context and trusting the result. A language model is only as good as the data you ground it in, and the retrieval, pipelines and infrastructure around it are where real engineering lives.
Articles in this hub
8 articles
IntermediateContext Engineering in 2026: What Replaced Prompt Engineering
Your prompt is fine. Your agent still loses the plot. Context engineering treats the context window as a finite budget — here's what goes in it, why bigger windows don't fix it, and the three techniques that do.
Read article
IntermediateQuantization Explained: How to Run a 70B Model on Consumer Hardware
A 70B model needs 140 GB of VRAM at full precision — until you quantize it. A practical guide to GGUF, K-quants, Q4 vs Q8, what quality you actually lose, and exactly how much VRAM you need.
Read article
BeginnerMCP Servers Explained: Build One, Then Run It Safely
An MCP server is how you give an AI agent real capabilities — safely. A practical 2026 guide: what the protocol actually is, how to build a server, what breaks in production, and the security rules you cannot skip.
Read article
BeginnerMachines Checking Machines: The Great AI-Detection Absurdity of 2026
Editors run AI detectors on writing, recruiters AI-screen AI-written résumés, universities falsely flag honest students, and paid 'humanizers' rewrite AI to fool AI detectors. An engineer's honest look at the absurd economy of machines checking machines — why AI-text detection is technically unreliable, who it harms, and what to measure instead.
Read article
BeginnerGenAI vs Agentic AI vs AI Agents vs LLM: What's the Actual Difference?
GenAI vs Agentic AI vs AI Agents vs LLM — the four terms everyone uses interchangeably, explained by an engineer. What each one actually is, how they nest, why GenAI and 'agentic' sit on different axes, with a comparison table and decision guide.
Read article
IntermediateAI Coding Agents in 2026: Claude Code vs Codex vs opencode
A vendor-neutral comparison of the three AI coding agents that matter in 2026 — Claude Code, Codex, and opencode: how the agent loop works, where each one fits, a decision table, and how to run them without handing over the keys.
Read article
IntermediateBrowser AI in 2026: Running Models On-Device with WebGPU and LiteRT.js
Running AI models directly in the browser is finally fast enough to be real. A practical 2026 guide to on-device inference with WebGPU and Google's new LiteRT.js runtime — what changed, how it works, and when to reach for it.
Read article
AdvancedReal-Time RAG in Python: Feed Your LLM Live Google Results (2026)
What if your RAG pipeline could pull fresh web context right before generating an answer? A step-by-step guide to building a live search retrieval layer with Bright Data's SERP API and Python.
Read article
FAQ
What is your AI engineering background?
Are you available to hire?
How do we start working together?
Taking AI from demo to production?
I build grounded RAG pipelines and reliable AI infrastructure — the data and MLOps layer that turns an impressive demo into a dependable system.
See AI engineering services →