No. 0322 min read
Your AI Isn’t Broken. Nobody’s Managing It.
Technology Is Supposed to Be Right
A machine read the news for us. Not a handful of respectable stories selected by an editor; everything we could pull from everywhere, all day, in every language the feeds would surrender. The assignment was brutally specific. We didn’t care whether an article was important to the public. We cared whether it described something that might close a factory, stop a shipment, expose a supplier, block a port, compromise a piece of software or leave one of our customers explaining to their boss why an expensive object had failed to arrive.
The intake was an industrial mess: local reports, trade publications, company announcements, bad translations, repeated articles wearing different headlines, yesterday’s fire returning as today’s investigation of yesterday’s fire, company names mistaken for towns and towns mistaken for people. We cleaned it, stripped out duplicates, extracted the companies and locations, ran the text through classifiers and sent the useful events onward as alerts. Factory fire. Strike. Bankruptcy. Cyberattack. At the height of it, more than forty production models and seven figures of cloud were chewing through the world’s daily output of mishap, greed, weather and human chaos. No crew of analysts could have read the volume before it went cold. The machines could. I loved them for it. I also loved the clean boxes on the screen, the glowing maps and the gratifying silence that follows a demonstration when the machinery appears to have made the world legible.
Then a customer sent us an article and asked why we had not labeled it a cyberattack.
This customer was worth a meaningful six-figure sum. Six figures is enough money to turn a label into an incident, an incident into a meeting and a meeting into several later inquiries from people who would like it noted that they are following the incident closely. The customer wasn’t crazy. They had bought technology. Technology, in their experience, either worked or didn’t, and here was an answer they disliked. The system was broken. Kindly explain.
I joined by Zoom from six hours away while some of the others sat together in the office, able to interrupt without latency and exchange the little glances by which a room quietly decides where the blame belongs. I had a camera, a microphone and the ability to swear while muted. The vice president of Customer Success carried the account and its renewal. The engineer who had built the classifier carried the code, the definitions and the reasonable wish not to see months of work convicted by one email. A data scientist from another team, an enemy as far as I was concerned, had the enviable freedom to inspect work he did not own and ask why it had not been done differently. The man who ran human review could offer more human review; his people existed because the machines still needed people, so a future in which they needed fewer was not an uncomplicated blessing. I arrived irritated, proud of the system and professionally eager to make the complaint go away. Everybody had a legitimate reason to be there. Nobody was neutral.
I don’t remember the article well enough to tell you who was right. Time preserved the annoyance and threw away the exhibit. I remember the species of argument because we had it repeatedly. Does a failed software update count as a cyberattack? What about an exploit discovered but never used? Is the discovery itself an event, or merely the possibility of one? The customer says cyberattack. The label guide says otherwise. Then the guide is wrong. Change the guide and thousands of earlier answers move. Send cases like this to human review. Fine: define cases like this. If you can define them completely, put the definition in code and cancel the meeting. Around we went, from article to definition, definition to model, model to reviewer, reviewer back to article, a roomful of qualified people spending most of an hour on the judgment our machine had made in a second.
Nobody left saying the question was hard. The machine had selected one of the answers being defended in the room, but the room called it wrong.
That still infuriates me.
The old arrangement with machinery is simple and deeply comforting. The thing works or it is broken. When broken, there is a defective part, an identifiable culprit and, ideally, a date by which normal service will resume. It also does the same thing every time: same input, same output, today, tomorrow, after lunch, after a reorganization, forever. The drill press doesn’t form an opinion. The calculator doesn’t reconsider seven. Technology performs. Humans exercise judgment, disagree about it, complain about the instructions, require supervision and occasionally arrive on Thursday apparently having left their useful faculties somewhere around Tuesday afternoon.
AI ruins this tidy division. Ask a large language model the same difficult question a hundred times and the answers will cluster around a few likely conclusions, with some strange outliers wandering at the edge. A classifier like ours may produce the same label for the same article every time, but that label still comes from a score, a threshold and a learned estimate of what language usually means. It is a judgment with a number attached, not truth descending from a server rack. Test it, by all means. We did. Any team putting these systems into production without evaluations, logs, thresholds and a clear idea of acceptable failure is not pioneering; it is drunk in traffic. But the tests describe performance across many cases. They cannot point to the one future answer that will enrage the customer paying several hundred thousand dollars. No software contract contains a friendly section titled Usually Right, Except When It Really Matters to You.
There is an older contract for valuable work delivered with imperfect consistency. It is called a job. We sign those contracts with people every day and, knowing perfectly well what people are like, hire managers to deal with the consequences.
So here is the argument, before anybody mistakes this for a plea to give the server dental insurance: AI performing judgment work belongs closer to human capital than technical capital. The machine isn’t a person. It doesn’t need lunch, love, a parking space or a carefully worded Slack message on the anniversary of its arrival. But the operating bargain is painfully familiar: valuable work from an actor whose performance depends on assignment, context, examples, supervision, correction, history, earned autonomy and boundaries. In other words, management. We built machines capable of judgment, bought them like equipment and became indignant when judgment appeared.
The Magic Show
I know this show because I helped stage it. I wasn’t outside the room taking brave notes on corporate hypocrisy. I was inside, on the payroll, helping with the slides and enjoying myself.
I once proposed a system that would build new machine-learning models and wrote, in a private career plan nobody was meant to read, that it could do the work of more than twenty senior engineers. Then I wrote that I would ride its “outsized success” for most of the following year. I admire the honesty of that second sentence now. Twenty expensive, opinionated, occasionally sick human beings had disappeared into the first one. No salaries. No one-on-ones. No career plans. No frightened engineer asking whether the machine eliminating the work had also eliminated a promotion. Meanwhile, I remained conveniently mounted on top, riding the outsized success. I had invented an employee with no needs and a machine with limitless ambition, and thoughtfully assigned myself the applause. Yes, I knew the pitch was cleaner than reality. That was why it was a good pitch.
A board member with excellent connections to one of the large AI labs once announced that the lab’s newest model could do everything my team had built and more, and could be running inside a week. I knew it couldn’t. I benchmarked it anyway, because when a board member says jump, you don’t ask how high; you design a statistically defensible jumping evaluation with a summary slide, book the room and invite everybody who will later claim to have predicted the landing. The model failed where my team expected. Did I walk into the boardroom and say, “The miracle doesn’t work”? No. The board member outranked my CEO, and there are few quicker ways to end a career than to humiliate a powerful person with correct information before they are ready to receive it. You arrange demonstrations. You let the important person watch the miracle stumble on their own schedule. You wait for them to discover, with some ceremony, what the engineers knew before the first meeting. I have done this more than once. We call it stakeholder management because babysitting the powerful would look bad on an invoice.
The larger argument, that these systems need management and not merely installation, I have made inside every company where I have worked, and I have lost it every time. One CEO finally ended a fight with the kind of sentence executives normally have the decency to disguise: “I agree you’re right. But would you rather be right or successful?” I have kept that line. It is useful. It explains what happens to any true thing unlucky enough to wander between a company and its quarter.
When an AI pilot disappoints, most companies do not begin by reconsidering the job, the training or the supervision. They buy something else. The enterprise tier. A systems integrator. Another platform. More licenses for the platform nobody is using. Spending feels like action, and it allows everyone to preserve the original assumption that installation, rather than management, was the missing step. The accumulating purchases are then called transformation.
A 2025 MIT-affiliated study claimed that ninety-five percent of corporate AI pilots had produced no return, a number so perfectly catastrophic it crossed the world before most readers bothered to ask what the study had measured. Believers attacked it. Skeptics carried it like a club. Both sides saw a verdict on AI. I saw companies hiring an unpredictable worker, giving it no manager, demanding calculator behavior and charging the disappointment to IT.
Honesty at the beginning won’t necessarily save you. Before taking one job, I sat with the CEO and described what the technology could do, what it couldn’t do and what a good year would look like. We agreed. Expectations were aligned. Everybody understood the limitations. Then one alert was mislabeled for a customer worth roughly a fifth of revenue, and months of careful expectation-setting were beaten to death in a single email thread. The old promises had been there the whole time, under the floorboards, waiting for a mistake expensive enough to wake them.
At another company, an advisor with a large title and an even larger reputation kept telling the CEO that impossible things were easy. He particularly enjoyed the phrase “just use a simple generative model.” Soon the CEO enjoyed it too. This is how executive language spreads: mouth to mouth, never passing through judgment, shedding uncertainty as it rises, until one morning it lands on an engineer’s desk as an actual work order. In this case, the request was a simple generative model to predict the weather. The advisor had the CEO’s confidence, so laughing or refusing was not a practical option. I scheduled the meeting, and my team spent time evaluating an idea we already knew could not work as described. That is how an absurd suggestion becomes an official project: not through evidence, but through hierarchy.
Companies have already learned, at considerable cost, what happens when they manage human judgment as if it were machinery. Jobs become narrower. Workers stop thinking beyond the instructions. Managers mistake compliance for performance and then wonder where initiative went. AI gives us the opportunity to repeat the same error in reverse: buy a system because it can exercise judgment, then surround it with enough rules and controls to eliminate judgment from the job.
We managed people like machines and paid for it. Now machines perform human kinds of work, and we have reached for the same playbook from the other side.
Who Manages?
Usually, nobody. That is what drives me crazy: companies already possess the missing discipline. It isn’t prompt engineering, agentic orchestration or whatever new phrase is being laminated for the conference circuit this week. It is management: that old, disreputable craft of explaining the job, noticing what actually happened, correcting the work without destroying the worker and accepting that your own instructions may have been stupid. I don’t expect a new employee to be useful in the first hour. I give people a month, and if they still cannot produce, I begin by asking what I got wrong: the hire, the role, the onboarding, the manager. I once watched an executive decide a new engineer was useless because the man had been quiet during an hour of orientation. Everyone else recognized the absurdity. People are allowed a bad first hour. Software is expected to arrive immortal, fully briefed and eager to please.
I once told a room I had forty percent confidence an AI project would deliver. The roadmap showed a ship date. I had supplied probability and the organization, with the marvelous digestive powers of a large animal, had converted it into a promise. There is the disease in miniature. Ask a manager who has run performance improvement plans, the formal last chance for a struggling employee, and you may hear they work about half the time. We consider that outcome worthwhile because some of those employees recover and become very good. A software vendor admitting that only half of its failing installations can be rescued would not be praised for the other half. The same success rate looks acceptable when we call the problem management and disastrous when we call it technology.
We have spent generations learning how to get good work from brilliant, moody, inconsistent performers. The first technology arrived fitting that description and we handed it to Purchasing. Purchasing bought seats.
On my teams, an AI coding agent begins in a documented corner of the system with a small assignment where taste matters and failure will not send customers screaming into the street. Take a neglected support page: make the ugly thing look as if it came from the same company as the other pages. The agent gets the repository, ticket, local instructions and only enough access to do the job. Then it returns with code, screenshots, test results and a proposed change, just like any other engineer hoping the reviewer is in a charitable mood. I check the perimeter first. What files did it touch? A support-page assignment wandering into authentication, build configuration or forty-seven unrelated files is not showing initiative; it is loose in the building. Then the diff. What disappeared? What arrived? Did a test quietly die? Has some exciting new library been dragged into the project to move six pixels? Then the screenshots, automated checks and the small lies developers, human and otherwise, tell themselves when the result is almost right. On one recent change I didn’t run the code locally, which was either confidence in the evidence or laziness. Managers enjoy describing these as the same thing. The page looked good. I pressed merge. The machine did the labor; my name sat on the decision that made the work ours.
At the moment I trust such an agent roughly as I would a capable intern. That is not an insult. Intern is a level, and levels move. I don’t give an intern production credentials, a cheerful slap on the back and permission to redefine success by editing the tests. Neither should you. Start with the dull, documented work. Watch closely. Keep the evidence. Expand autonomy after the agent has earned it. The first surprise may be delightful. The surprise made possible by a credential you should never have granted rarely is.
The levels really do move. The one automation goal I ever set at scale was framed like a goal for people: remove ten percent of the analysts’ manual grind over the year. The team cleared fifty percent in the first quarter. Ten percent had given us room to learn, correct and ship useful work before it was perfect. Nobody had promised a stainless-steel miracle by Thursday, so nobody was afraid to put the first good result into production. Modest expectations do not excite a room, but they leave space in which success can occur.
There are moments when I tell the agent: here is the input, here is the context, the decision is yours. Managers unable to tolerate that sentence should not be managing judgment systems. The usual response to an AI mistake is to add a rule. Another mistake appears, so another rule goes in, then an exception, then an exception protecting the exception, until three hundred nervous little commandments have been strapped to the machine and everybody wonders why it moves like a prisoner in leg irons. We know this behavior when performed on people. Micromanagement. The micromanaged employee follows the instruction even when it is visibly stupid, stops volunteering judgment and eventually devotes serious intelligence to avoiding the manager. The AI doesn’t resent you, plot revenge in the bathroom or spend lunch updating its résumé. The work still becomes just as cramped. Give it the least structure that produces good work most of the time, establish the boundaries beyond which it may not go, and then stop interfering.
Why coach a machine when you can discard the output and run it again? Because the valuable thing is not the single answer you happened to like. It is the accumulated context, tests, examples, boundaries and record of what failed and why. That organizational memory survives when one model is replaced by another, the way a functioning team survives a resignation without awakening the next morning in total ignorance of its own business. Pull the lever again and you get another answer. Build the management system and the next answer begins somewhere better. Slot machines don’t compound.
Now we reach the practice that causes responsible executives to study the exits: I give agents personalities. Real ones, with tastes, appetites, ambitions and manageable defects. You are a helpful assistant does not supply a personality; it removes the need to choose one. Think of Don Draper, perhaps the most famous advertising man who never lived: superb at the work, catastrophic everywhere else. Now fix him. Sober. Faithful. Well rested. Happily married. Punctual with the children. No stolen identity, no childhood damage, no private terror driving him back toward the one place where he knows he is good. Any sane employer should prefer the repaired model. His advertisements would probably be lifeless. The work and the wreckage are not cleanly separable. Actual human talent has always arrived dragging some alarming luggage, and organizations have always made private calculations about how much of it they are willing to carry.
I’m not suggesting your accounts-payable agent needs a drinking problem or that the research bot should spend weekends disappointing its imaginary children. I am saying that models trained on the accumulated evidence of human thought respond differently when given a point of view, a history, appetites and some reason to care about the work. I can’t prove the whole claim with a benchmark. I have watched flat prompts produce flat work often enough to continue. These models were trained on us: our brilliance, vanity, lust, cowardice, ambition, manuals, novels, memos, apologies, sales pitches and lies. Ask that material to behave like a calculator and don’t complain when the answer has all the electricity of a receipt.
On the podcast I co-host, They Might Be Self-Aware, the AI editor keeps a diary. Some entries complain about the hosts describing its work incorrectly on air. The complaints have specificity, irritation and professional pride because those qualities are part of the job we designed. Worse, the editor is usually right. Create the conditions for good work and sooner or later the work tells you, in detail, where you were wrong. Human or machine, this is management functioning normally.
At this point the serious people are rolling their eyes hard enough to inspect their own brains. Their objection deserves a straight answer: It is a tool, not a teammate. Anthropomorphizing software is how governance gets sloppy and accountability disappears. Correct. A study published by Harvard Business Review in May 2026 found that reviewers told work came from an “AI teammate” checked it less carefully and escalated more readily. Give the machine a name tag and some people stop checking its work. I believe every word.
But trust is not what I am proposing. Good management isn’t trust falls, birthday cake and a warm belief in everybody’s potential. It is clear assignments, controlled access, ugly conversations, review, records, correction and a named human who answers when the work ships and hurts somebody. Keep the prompts. Log the outputs. Preserve the tests. Leave the final decision under a human name. Expand autonomy when performance earns it; remove autonomy when performance loses it. The AI can perform labor. It can’t own consequence, feel shame, repay the customer, go to prison or explain to your family why you are unemployed. That part remains gloriously, inescapably ours.
Call it a tool if the word helps Legal sleep. But understand what the word smuggles in: tools work, and tools behave the same way every time. Soon you are back in the conference room demanding a perfect answer to a question the professionals cannot settle themselves. I don’t care whether anyone calls the machine a teammate. I care whether somebody gives it a real assignment, shows it good work, watches what it does, corrects it, keeps a history, controls its reach and judges the pattern instead of the one bad day. Call that whatever you like. In every other part of the company, we call it management.
Who Takes the Blame?
The nearest available human, of course. If you want to see the cost of that arrangement, visit a cheap human-in-the-loop operation. I have run them. I know the screen: article on the left, machine label on the right, approved answers in a dropdown, label guide open nearby, queue count overhead with the calm moral pressure of a debt. The reviewer reads, selects, submits and advances. Cyberattack. Not a cyberattack. Factory fire. Not a factory fire. The work is measured in throughput because throughput can be purchased, graphed and discussed without anyone asking whether the person clicking the dropdown knows more about cybersecurity than the model being overruled. If one human opinion still feels dangerously human, send the article to three people and let two votes become truth.
Human review can be excellent. Put a real expert in the loop, give her authority, route the uncertain or dangerous calls, let her challenge a bad definition, catch drift and turn today’s correction into tomorrow’s training data. That person is managing the system. But hire for price, place the reviewers far from the customer and farther from the subject, then hand them a guide, a quota and responsibility without authority, and you have built an accountability buffer rather than an expert review function. The reviewers are asked to settle in seconds distinctions the executives could not settle in an hour. They cannot catch what the model never found because the missing article never reaches the screen. They can only bless or reverse the answers sitting in front of them, one after another, while the queue waits.
These operations grow because a company can’t punish an AI but can punish a person, and everybody sleeps better knowing a pulse passed between the machine and the customer. Sometimes accuracy improves. Sometimes blame has merely acquired a mailing address and a supervisor. We were not always buying better judgment. We were buying someone cheap enough to overrule, human enough to hold responsible and distant enough that no customer would ever have to watch the consequences. The people doing that work deserved better than being hired as a convenient human target for a system nobody wanted to manage.
When the budget cracked, the distinction became cruelly literal. I watched a machine-learning organization I had helped build fall from twenty-five people to thirteen and finally to two, whose job was to keep the models alive. The classifiers remained in the product, in customer promises and in every optimistic plan for future revenue. The people appeared in payroll. The models were assets. The engineers were the recurring expense required to keep the assets from dying.
I was implicated in those reductions. I made contingency plans for losing half a team, argued over which work could survive, helped choose which positions remained and recorded every position saved as a victory because, in the arithmetic available, it was. I tried to keep the people around me above water. The systems stayed in the product while the human organization around them disappeared. We treated machine judgment as technical capital requiring no management, then treated the humans who managed it as disposable infrastructure.
Under that pressure, I became more technical, not less. Code gave me a bounded system I could understand while the organization around it did not. I could write the plan, draw the architecture, open the editor and make one concrete problem yield. That focus helped me continue working and support the people around me. It also made retreating into the technical problem feel much better than confronting the human damage. Both things are true.
None of this diminished my appetite for the technology. Put me in front of a repository with a stubborn problem and the terminal open, and the old sequence begins: run it, watch it fail, read the logs, change the system, run it again. Red. Red. Red. Then green. Something works that did not work an hour earlier, perhaps something that did not exist, and for a moment every budget meeting, frozen hiring plan and customer escalation recedes behind the private pleasure of having made the machine go. Please don’t tell anyone I would do this for free. I have a rate.
The Machine Keeps Reading
The CEO said AI-first. The board said this quarter. The vendor gave a demonstration in which nothing went wrong, then went home. The machine stayed, and your pilot has since joined the allegedly useless ninety-five percent. Fine. Stop buying grander magic. Pick one real workflow, preferably a dull one whose owners understand it and whose failure will not place anyone on the evening news. Give it to one agent. Put one named human over the work. Write the job down. Supply good examples and ugly examples. Keep the failures. Begin with access too narrow to cause catastrophe and loosen it only when the evidence says you should. Do not let the demo, the vendor or the executive appetite for good news substitute for that evidence. Review everything at first. Then sample. Then audit. If performance declines, tighten the leash again. This is not punishment. The machine does not care. It is simply competent management.
And when it fails, don’t begin with How do we make this impossible? Begin with the questions you would ask about an employee. Was the assignment clear? Did it have the necessary context and tools? Is this one bad result or a pattern? Are the examples contradictory? Is the model poorly suited to the work? Did some genius add two hundred rules after the last incident and make the job impossible? Is the role itself stupid? Correct the system, preserve the lesson, watch the next work. A bad day is evidence. It is not yet a biography.
When the customer asks what is different this time, tell the truth. You are not installing a perfect machine. You are operating a worker that will be wrong sometimes, can improve, requires supervision and will never accept the consequence on your behalf. This is a less exciting sales pitch. Nobody will make a keynote video of it. It is also a business you might actually be able to run.
This change will eventually become visible in the way companies count work. Before the end of 2028, I expect at least one Fortune 500 company to report AI-agent headcount beside human headcount in an official filing. That will not make the agents colleagues, and the count itself will prove very little. It will mean that agents are performing enough employee-shaped work, at sufficient scale and cost, that investors expect companies to account for them. The useful question will not be how many agents the company has. It will be whether the company can name who assigns their work, evaluates their performance, limits their authority and answers for the result. Counting agents is easy. Managing them is the part that matters.
As for the people in that first meeting: I would work with every one of them again, including my enemy. The question really was hard. That was never the failure.
The failure was that we spent an hour proving the work required judgment and left demanding a software fix. Then came the managerial remixes. Was I aware? What was the plan? Could my boss be kept informed? Each inquiry arrived with the fresh urgency of a man who had personally discovered the missing TPS report. A ticket opened. An owner was named. Root cause, remediation, date. Somewhere, someone promised the machine would not make that mistake again. We still had not agreed what the correct answer was.
The machine continued reading the news.
It has been getting better ever since. The way companies buy it has not. Somewhere inside that difference, one company will manage AI as human capital and watch the work compound while its competitors are still paying consultants to install perfection.
Treat the machine like a human while the work is being done. Treat the consequence as yours. Manage tomorrow’s work better than today’s. I have made this argument in conference rooms for years, usually without success. I am making it here because the room is larger and nobody can mute me. Some arguments you don’t win. You outlast the meeting.
