Inside OpenAI’s Race to Reinvent Software Development for the Agent Era by Laura Entis
A few interesting things in this article including the art and visualisation.
The first line of defense is continuous integration (CI), the automated checkpoint that tests new code before it merges into a shared codebase. Farther along is continuous deployment (CD), which ships approved code into production. Built in a pre-agentic world, both were designed for code written at a human pace. But as the cost of generating code falls toward zero and volume explodes, the number of changes waiting to be tested and merged will overwhelm the system. Problems CI fails to catch pass through CD into production, where failures become incidents.
The solution, in Venkataramani’s view, isn’t simply to expand the existing system’s capacity. “If you just scale the CI infra and the CD infra, you’re just going to have 100x more code going in every day,” he says. “And then what?” Scaling CI and CD moves more code through the system without making that code safer.
His solution resembles modern stormwater management. Instead of relying on a single downstream checkpoint, it filters for different problems at multiple points. “You add 10 such filters,” he says, “you’re going to eventually have clean water.”
About adding multiple filters or agents with skills that check for specific things.
And further on, about how the industry may look like in the future.
As the models improve, Venkataramani expects code review to give way to prompt review, and then plan review. Engineers will no longer review lines of code and instead evaluate the intentions behind them, starting with specifications, and then architecture documents, and eventually the business problem itself. New coding environments might return only a high-level explanation of what the code does, he says. A complex project that today requires millions of lines of code could be expressed as a short set of instructions laying out what the system should do, with the model filling in the rest.
This, unsurprisingly, changes the job requirement. Instead of solid executors, you need people who can answer messy, ambiguous questions that don’t have clear-cut answers, like how to determine whether to fund a project or the best way to measure whether an initiative is paying off. OpenAI has long hired people who are “very entrepreneurial, very self-directed, who have their own opinion and really strong ownership,” Tang says, but these qualities have calcified into non-negotiables as the models absorb more of the operational work. “We only hire those people now.”
None of this was entirely new. But I always find it interesting when names and pictures are assigned to engineers and these were all infra people interviewed for this piece. So I could relate more.
On AI Coding and Its Discontents - Cal Newport
“Writing your own code, slowly but surely, and using LLMs for narrow or particularly annoying tasks (say like writing tests or throw-away scripts), is the best way to produce the highest quality code, since it’s the only way to properly understand it.” […]
This last year has been exhausting. The PR departments of the frontier labs have done an excellent job convincing us that AI developments are occurring at an astounding, world-changing rate. But if you zoom out, it becomes clear that almost every “breakthrough” since last summer has concerned the narrow domains of computer code and math, which are defined by highly structured languages and come accompanied by massive amounts of specialized training data.
And yet, even in this best-case-scenario setting for AI, we’re still struggling to figure out how to actually use these tools in a way that makes sense in the long run.
Writing production code is tricky in that it can not fail. In this engineer’s case it broke production twice.
The thesis of this article is that it’s ok to use AI to write code, which if it breaks does not cause to too much of an issue. Which has been my experience too. I have primarily used Claude to build personal tools, and my website and some scripts and what nots at work. Also, how much of code out in the world is truly code that can not break?
But I know of people who are talking to their agents and having it build things.
The one thing I agree with is - that we are all trying to figure out the steady state of using AI in development work.
Give me one reason by Seth Godin
The honest rejections would say something like, “my boss didn’t like you,” or “I was in a bad mood,” or “we realized that our spec wasn’t very clear and the person we liked didn’t fit it either,” or “we wanted someone who looked more like us,” or “your competence made us nervous,” or…
Take what they wrote and replace it with “no.”
It’s a short read. And I just wanted to share it.
Defending my own brain against enshittification
His days are punctuated by cups of coffee and tea at strict intervals, cigarettes, ironed sweater vests, empty time for deep thinking, placing a bench outside his office for passersby each day, and most notably, when he arrives at his office at 10 am alone, he opens all the windows in the building and says good morning to the ghosts who live there.
Lovely story.
I’m 38 and I Can’t Support Myself Anymore by vōx
When you become chronically ill or disabled, you begin to understand how deeply this ideology has embedded itself into the psyche. The shame of no longer being able to work consistently feels personal. As though my body’s inability to keep pace with capitalism says something fundamental about who I am.
This doesn’t mean you should never grind at 100% effort. I think there are probably two or three times a year where I work as hard as I possibly can: long hours, intense focus, thinking about the problem from when I wake up to when I go to bed. But I reserve this mode of work for when the rewards are really high. For the rest of the year, I take it relatively easy.
There is a matter of luck involved in this too, or skill. And the headline is so damn salicious!
Vacations and Noticing Things
Wheezing past the cradle
The Human Cost of 10x AI Productivity by Denis Stetskov
They found three mechanisms of “workload creep.” Task expansion: everyone’s scope inflates because AI makes it possible to do more. Blurred boundaries: AI prompting happens during lunch, commute, evenings. Implicit pressure: when colleagues visibly do more with AI, expectations rise for everyone.
[…]
The mechanism is asymmetric. When I write code, I externalize a mental model that already exists. The thinking is done before the typing starts. When I review AI-generated code, I have to reverse-engineer somebody else’s reasoning out of an artifact produced by a system that has no idea what our business does. Fundamentally harder.
Code is run more than read by Facundo Olano
This phrase is, by now, common programmer knowledge, a reminder that the person first writing a piece of code shouldn’t buy convenience at the expense of the people who will have to read it and modify it in the future. More generally, code is read more than written conveys that it’s usually a good investment to make the code maintainable by keeping it simple, writing tests and documentation, etc. It’s about having perspective over the software development cycle
There is a nice progression in this article. The final stage is - biz > user > ops > dev.
Protect Your Shed by Dylan Butler
Earlier in my career, I was new to containerisation and cloud infrastructure, and the learning curve at work was steep. But because I was standing up containerised systems and running them on GCP at home on my own time, the concepts landed faster. I was getting reps in from both sides.
A metaphor.
I Deliver Parcels in Beijing
The Only Moat Left Is Money - Elliot Bonneville by
Reach is also gravitational. Past some threshold it accumulates without you — posts find people, people find posts, the thing feeds itself. Below the threshold, identical effort produces nothing. Same quality, same idea, same work. Zero. Not because it was bad. Because you showed up on the wrong side of the line.
Stop generating, start thinking - localghost
In the wake of the Horizon scandal, where innocent Post Office staff went to prison because of bugs in Post Office software that led management to think they’d been stealing money, we need to be thinking about our software more than ever: we need accountability in our software.
Nice write up. I see the same points that I’ve seen elsewhere. I don’t think it’s the engineers driving this revolution though. It’s the business leaders doing that. There is a fomo in the industry - people are committed to AI without having any use case for it.
Regarding accountability it will fall on the reverse-centaurs, the people left to deal with the large amounts of AI generated work. Nobody will say Claude made a mistake, it’s the employees who made mistakes by not verifying what was generated. It will not be a great place, but we are barreling toward it.
Be Wary of Digital Deskilling - Cal Newport
In his 1974 book, Labor and Monopoly Capital, the influential Marxist political economist Harry Braverman argued that the expanding “science-technical revolution” was being exploited by companies to increasingly “deskill” workers; to leave them in “ignorance, incapacity, and thus in fitness for machine servitude.” The more employees outsource skilled activity to machines, the more controllable they become. […]
Boris Cherny is a senior technical lead at Anthropic who manages a large team and likely owns a significant amount of stock options in the company. Of course, he’s excited about the idea of agents replacing programmers, but that doesn’t mean we have to share his enthusiasm.
Creativity Inc
'The Downside To Using AI for All Those Boring Tasks at Work' - Slashdot
Roger Kirkness, CEO of 14-person software startup Convictional, noticed that after AI took the scut work off his team's plates, their days became consumed by intensive thinking, and they were mentally exhausted and unproductive by Friday. The company transitioned to a four-day workweek; the same amount of work gets done, Kirkness says. The underlying problem, according to Boston College economist and sociologist Juliet Schor, is that businesses tend to simply reallocate the time AI saves. Workers who once mentally downshifted for tasks like data entry are now expected to maintain intense focus through longer stretches of data analysis.
This is an interesting problem. I see little discussion of it elsewhere. What will happen? Will we continue to work the same hours doing more, or will we be working less doing the same.
Knowledge work is highly cerebral in nature. That requires down time, in order to continue working at a high level.
Remote Work is Officially Dead, Says the World's Largest Recruiter - Slashdot
"You have to be very special to be able to demand a 100% remote job," van 't Noordende told Fortune. "That's increasingly the story. You have to have very special technology skills or some expertise." The equilibrium appears to be settling at a hybrid model of three to four days in office for most workers.
That has been my experience too.
Also even before the Covid pandemic, the fully remote option was there for high performers or edge cases, where people had specific requirements to work from home and were good enough that they could not be kicked out of the job.
Transparent Leadership Beats Servant Leadership by kqr
The middle manager that doesn’t perform any useful work is a fun stereotype, but I also think it’s a good target to aim for. The difference lies in what to do once one has rendered oneself redundant. A common response is to invent new work, ask for status reports, and add bureaucracy.
A better response is to go back to working on technical problems. This keeps the manager’s skills fresh and gets them more respect from their reports. The manager should turn into a high-powered spare worker, rather than a paper-shuffler.
Interesting comparison between parenting and managing people here.
Claude Opus 4.5, and why evaluating new LLMs is increasingly difficult by Simon Willison
Anthropic released Claude Opus 4.5 this morning, which they call "best model in the world for coding, agents, and computer use". This is their attempt to retake the crown for best coding model after significant challenges from OpenAI's GPT-5.1-Codex-Max and Google's Gemini 3, both released within the past week!
I did not have preview access to Opus4.5. Nor do I need it for the things I generally use LLMs for.
With the base text only models, I guess there is no more step change now. They may show benchmarks that they are the best model for coding, but it’s single decimal points. It does not really matter.
What matters more is the features they add - like when Anthropic added the skills feature. What you can do is more important. And yes I still believe it will be human in the loop situation. Will we be centaurs of reverse-centaurs is an open question.
Work at a Natural Pace
Obsess Over Quality
Do Fewer Things
Code like a surgeon by Geoffrey Litt
A lot of the “secondary” tasks are “grunt work”, not the most intellectually fulfilling or creative part of the work. I have a strong preference for teams where everyone shares the grunt work; I hate the idea of giving all the grunt work to some lower-status members of the team. Yes, junior members will often have more grunt work, but they should also be given many interesting tasks to help them grow.
With AI this concern completely disappears! Now I can happily delegate pure grunt work. And the 24/7 availability is a big deal. I would never call a human intern at 11pm and tell them to have a research report on some code ready by 7am… but here I am, commanding my agent to do just that!
The idea being AI works on the secondary stuff and keep it ready while you work on the primary stuff.
I found the above idea important as well, to rotate grunt work among the full team. I have had this in the past where senior members would not work on tickets, etc.
We try to make sure everyone works on everything.
Better Than Average
What to Automate
A Little Inefficiency Is Good
How to Handle Stress at Work
How to Work With Your Boss
Majority of the Organisations Are Not Seeing Any Monetary Benefits From Deploying AI
Art is a project. Connection, community building, counseling–all of these are projects. When our work is project-focused, we’re not a cog in a vast machine. Instead, we’re a contributor with agency, someone who is working with and for the agenda we’ve agreed to.
The Bad bosses try to have it both ways. They are stingy with agency, authority and compensation, and insatiable when it comes to effort. But smart leaders understand that given the chance, most of us would love the chance to be seen, to contribute and to be part of something.
Be a Hybrid
Have expertise in two or more things
About Reflections on Writing
From people who have been doing this for many years
Why Work?
The myths of work
The Last Work Left in This World
Train the models!
How to Complain
Or, how to make your boss's life easier
About Glue Work
Is glue work bad? Depends.
Two Lessons on Work
Show your work + Ask for help
Types of Workers in an Organisation
Or, evolution of the type of worker you are
Trusting People to Do the Work They Were Hired to Do
Curb micro-management