- AI engineering
- careers
- agents
- evals
The moat was never the code. It was everything around it.
AI writes most of my code now. What's left for a mid-career engineer: evals, agent architecture, context engineering, AI governance and domain depth.
Published: ·5 min read·First on Substack
YKYasir Khalid
For the past few months I've been having a quiet crisis at work. I'm a senior developer at HSBC, optimising legacy systems, orchestrating AI agents that talk to each other, helping relationship managers do their jobs faster. It's genuinely interesting work. But somewhere along the way I noticed something unsettling: I barely write code anymore. Most of what I produce is AI-generated. My job has become prompting my way through problems. And I started asking myself - what exactly is the value I add here?
I don't think I'm alone in this. A lot of engineers in their mid-career right now are sitting with the same discomfort, and most of the advice they get is generic. Learn Python. Learn Machine learning (drilldown on basics). Stay curious… None of that addresses the real question, which is: if an LLM can write better code than most principal engineers within the next few years, what does a human engineer actually bring to the table?
The honest answer I've landed on is this: coding was never the moat. It was the bottleneck. The actual value was always judgment, domain knowledge, and systems thinking - but those were hard to see because they were wrapped up in the act of writing code. Now that AI has removed the coding bottleneck, everyone is being forced to compete on the stuff that was always more important and always harder to fake.
The specific skills that matter right now - at least for someone in my position - aren't the ones I would have instinctively reached for. When I brainstormed what to learn next, my list looked like this: linear algebra, differential equations, Machine learning concepts, GoLang for concurrency, advanced SQL optimisation, blockchain, UI/UX design, quantum computing (I’ve shown this list to a few friends of mine to get their thoughts on this and what they think I should’ve picked next). It felt like a serious list. It was almost entirely wrong.
Linear algebra and differential equations matter if you're building models from scratch. I'm not. I'm building systems on top of models - which means understanding that attention mechanisms use matrix multiplication gives me almost nothing practically useful. GoLang is genuinely good for high-performance concurrent services, but I already have Python and limited Scala; adding a third language before consolidating the skills that actually matter is just spreading thin. Blockchain has been the next big thing for a decade and produced very few production use cases outside of cryptocurrency. Advanced SQL optimisation is valuable for a specific DBA-adjacent role, not for where I'm heading. Quantum computing is fascinating and irrelevant to my career for at least the next ten years. The whole list was a classic response to career anxiety - casting wide, grabbing things that sound impressive or fundamental, mistaking breadth for progress.
- Linear algebra
- Differential equations
- Machine learning concepts
- GoLang for concurrency
- Advanced SQL optimisation
- Blockchain
- UI/UX design
- Quantum computing
- Evaluation and observability
- Agent architecture
- Context engineering
- AI security and governance
- Domain depth
What actually matters, I think, breaks down into a few things that are much less glamorous but much more valuable.
The first is evaluation and observability. In a regulated bank, "it seems to work" is not a position you can hold. If an AI agent is helping 15000 relationship managers make decisions about clients across a hundred geographies, you need a systematic way to know whether it's getting better or worse over time, what it costs per query, where it fails, and how you catch regressions before they become incidents. I already use Langfuse for tracing, but there's a significant difference between having used something and owning the discipline. The person who can walk into a senior meeting and show a quality trend dashboard, an LLM-as-judge scoring system, and a prompt regression suite that runs automatically on every change - that person is the most credible voice in the room on AI. Almost no one is doing this rigorously right now.
The second is agent architecture. I work with Google ADK, which abstracts a lot away. That abstraction is useful until it isn't - until something fails in production in a way the framework didn't anticipate, or you need to make a principled decision about whether to use a single agent or a multi-agent pattern, or you need to explain to a sceptical principal engineer why you made the choices you made. Understanding the underlying patterns - orchestrator-worker, parallelisation, evaluator-optimiser, human-in-the-loop checkpoints; and more importantly understanding failure modes, means you're not locked to one framework and you're not guessing. My background in distributed systems is actually directly useful here: an LLM tool call is just a non-deterministic network call. Treat it like one. Timeouts, retries, dead-letter queues, fallbacks. The patterns transfer.
The third is context engineering, which sounds like a buzzword but is actually the most concrete skill in the list. The model's raw capability is fixed. What you put in front of it determines almost everything else. How you structure a system prompt, how you design tool schemas so the model calls them reliably, how you handle retrieval so the right information is in context at the right time, how you manage a conversation that's running long; these are real craft decisions that have a direct impact on whether an agent works in production or just in demos.
The fourth, and probably the most underappreciated, is AI security and governance - especially in financial services. Prompt injection attacks, where malicious content in a tool response tells your agent to do something it shouldn't. Agent sandboxing, defining which tools require human approval and which actions are irreversible. Audit logging that captures not just what an agent did but why, so that when a regulator asks you to explain a decision the agent made six months ago, you have an answer. These aren't hypothetical concerns. They're regulatory requirements waiting to happen, and the person who's already thought them through carefully is in a very different position than the person who hasn't.
On the question of system design and judgment - which is ultimately what all of this rests on: I've stopped believing that courses or books are the primary path. Designing Data-Intensive Applications is worth reading, not to memorise it but to extract mental models: the chapter on replication tradeoffs, the one on consistency versus availability. These give you vocabulary for decisions you're already making intuitively. But the actual development of judgment comes from putting yourself in decision-making situations faster than your years of experience would naturally allow.
There's one more thing I keep coming back to, which doesn't fit neatly into a skill category. I work in financial services. I've been here a few years but I'm still scratching the surface of how correspondent banking actually works, what compliance officers are genuinely worried about, what "know your customer" means at a technical implementation level. This domain depth is something an AI genuinely cannot replicate, and it makes every architectural decision sharper because you understand the constraints you're operating in. An agent that works technically but creates a compliance problem isn't a win. The person who understands both sides of that sentence is the person who's hard to replace.
The summary version is this: AI gets better at writing code every month.
The moat was never the code. It was everything around it.