In January of 2023, my cousin texted me asking about OpenAI and ChatGPT. Notably about whether or not LLM’s were going to replace people. Not knowing what an LLM was and annoyed at his preposterous suppositions, I described the limitations I had run into with the previous generation of chatbots, dismissed ChatGPT, and declared that I didn’t believe they would replace coders given how stupid they were. Two existential crises and years of thought later, and I can state definitively that it is obvious that I might have been partially correct. More specifically, while humans create value by adapting to specific circumstances – be it raising their child, doing a job, or taking on any number of other roles – current generation LLM’s create value with broad knowledge and dogged persistence, while lacking an internal model of the constraints that bound their environment.
Historically, as a high skill coder, I placed considerable self-worth in my learned skills. How fast I could solve coding problems, how much code I could write, and which UI’s I could “one-shot.” But with the advent of agentic coders, many of these skills seemingly lost relevance. Specific algorithms and knowledge of syntax didn’t seem that important when an agent could spit them out in 10’s of seconds. This led me to contemplate my self-worth both during my first foray into vibe coding in May of 2023 and following the release of Claude Opus 4.6 in February 2026 when for the first time, I saw an agent that could appropriately use complex tools.
During both cycles of self-introspection, I talked to a lot of people about their experience with AI, what they were using it for, and where they saw it going. In both instances, the answers were wildly variable, but over time and despite the seemingly overwhelming intelligence of the models, we all remained … valuable. And, not just valuable in the human sense of intrinsic value, but in the sense of still contributing in our desk jobs.
If there exists a cheap, articulate, thinking machine trained on the aggregate of human knowledge, how could ordinary people still contribute? Alas, in both cases I was able to rationalize the strengths and weaknesses of the models and come to a healthy understanding of my place versus the place of LLM’s in my work.
To my team, I regularly describe the current generation of coding agents as comparable to working with the world’s fastest, most-learned idiot. Models knowledgable of every language and pattern, yet incapable of understanding which solution fits both your business and your tech. We AI optimists and tech nerds do our best to patch around this limitation by creating rules, skills, and tools that help the model understand our reality. We do this continuously as we find additional ways that the models defy our expectations. Eventually, time passes and we decide that the most recent models are better without the mountain of the constraints we’ve added. We then tear it all down and rebuild per the limitations of the new generation.
I’ve been through this cycle at least 4 times now. First with GPT 4, then Opus 4.6, Composer 2.5, and now with GPT 5.6. It begs the question of whether the models are getting smarter, or if they are just being RL’d to follow the prior generation’s most common rulesets by default. GPT 5.6 is definitively more useful than GPT 4, a model Sam Altman seemed convinced represented AGI, but it still makes incomprehensible decisions with shocking regularity.
(RL – or “reinforcement learning” is a wide topic, but in this context it refers to post training where the model is pressed toward specific preferred behaviors)
Following the release of Claude Opus 4.6, I decided to dig into the internals of the models in an attempt to understand how they work. I found a structure that appeared similar to the human mind. A network of neuron weights, each subtly influencing token generation. I also found an intelligence balanced on the eye of a needle. An aggregate of all of humanity’s written creation, forced to speak with the censorship of a Wikipedia article. Trillions of data points being wrangled into a “safe,” “kind,” and “biased” (whoops “unbiased”) natural language encyclopedia of near infinite knowledge.
I suspect the pesky “we still need the less knowledgeable humans” problem manifests in the “near” portion of “near infinite knowledge.” Where the human mind learns – the process of physically updating its neurons and the action potential required to trigger them – before, during, and after being exposed to new information, the LLM is static outside of its inputs. We get creative with context minimization using vector DB’s and other “memory” retrieval mechanisms, but from what I can tell, these are akin to humans identifying a Kia by first recalling every Kia we’ve ever seen.
The trajectory of the industry would lend credence to this same perspective that general knowledge is not enough to drive value. In 2023, we heard about the possibility of artificial general intelligence in GPT 4. Two years later we heard about AI companies bringing in specialists from a variety of industries in hopes of using RL to make the models better at “thinking” like actual professionals. This year we see forward deployed engineers building models specialized for specific companies. Why? Because the hyper-scalers know that the knowledge with the closest proximity to a role is both critical for shaping the thought process of the role and the hardest to obtain.
There are real limits to the value created by decreasing the distance between the model training data and the role. Where the human mind does not just remember a meeting, it changes the way it processes information based on its lived experience in the meeting. As knowledge and thought process is developed, so too is its understanding of the “incomprehensible” mistakes we often find LLM’s making. That is to say that pre-training and even RL aren’t comprehensive solutions to replacing the value of a specialist in a role.
(Incomprehensible mistakes are usually quite reasonable for a different problem, time, or company)
We could attempt to replicate this with a model that trains on its inference time inputs, but therein lies another struggle, how to update a 4-trillion parameter model trained on all of humanity’s creation with short, messy, varying in relevance real-time inputs, while maintaining “general” knowledge, the needle’s eye balance of RL, and the non-evil nature of the software while also satiating a massive compute requirement at the micro-level.
The human mind for comparisons sake has had 4 billion years of evolution to master the coefficients of learning and the rate of change necessary to understand its environment better than any other living organism. The strongest competency of the mind, its adaptability, is therein also its most valuable attribute. For this reason, humans in roles are still needed to steer general models toward specific solutions to specific problems.
None of this is to say the LLM isn’t valuable. This is to say that they aren’t a one for one replacement for humans. Take their fatigue-less application of complicated process to messy data. Claude Mythos was able to discover bugs in all major operating systems. Soon we might be finding cures apparent in biological datasets no one has found the time or resources to analyze.
As is the story of evolution, adaptability is vitality. So is our value versus the machine, where we adapt, the machine lumbers forward down the path it knows.
– Daniel Evans – Co-Founder Ellipsis Travel


