Stop Chatting With AI — Start Managing It Like Staff

Most people are using AI wrong, and it isn’t because they picked the wrong model.

They open a chat window, type a sentence and a half, get something mediocre back, and conclude the tool is overhyped. What they’ve actually done is hire the smartest employee they’ve ever had access to and then give that person no job description, no SOP, no deadline, and no way to check whether the work got done.

I spent a session with Joe and Sean working through the alternative. Not theory. Two prompts, dictated out loud, live, in front of them. Here’s what that looked like and why it works.

Your prompt is too short because your typing is too slow

The first thing I did was turn on voice dictation. Not because talking is trendy, but because typing is a bottleneck on how much context you can afford to give.

When you type, you unconsciously compress. You write “help me improve my ops” because writing the full picture would take fifteen minutes. When you talk, you’ll give three minutes of context without thinking about it, and three minutes of speech is roughly five hundred words of specifics the model would otherwise have to guess at.

Here’s the important part: this didn’t work three months ago. Dump a wall of text into an older model and it got confused, latched onto the wrong sentence, and gave you something worse than a short prompt would have. The models are now strong enough to actually process all of it. So the old advice to keep prompts tight is out of date, and most people haven’t updated.

The prompt I dictated asked for a fifteen to twenty page PDF that was honest and truth-seeking. I said explicitly: don’t be afraid to say things that might be embarrassing for me, because the goal is to improve operations, and that means improving me. I asked it to map how I’m using it now, where my direction is vague, and what my job function actually breaks down into.

Screenshot 6

Then the part most people skip. I told it what to do if it didn’t know something. If you don’t understand my role, take your best guess, then come back and have a conversation with me until we get it right. That single instruction turns a one-shot request into an iterative working relationship.

Operate as a department, not a person

The goal I gave it was specific: help me operate as a department instead of a person.

That means my job gets decomposed into logical sub-tasks that run on a repeatable basis, each with a clear way to determine whether it was done successfully, so a sub-agent can execute it and another agent can measure it. I’m not the one doing the work anymore. I’m the department head managing agents who are equipped with skills and running as scheduled jobs.

Ten times the impact doesn’t come from working harder. It comes from converting the things you do repeatedly into things that don’t require you.

The second prompt: build the fleet

The second prompt was blunter. I told it I’d been mainly chatting and not unlocking the power of agents, and that agents are coworkers I manage who have clear job descriptions and follow published SOPs that get constantly updated.

Then: set up twenty-plus scheduled tasks that run daily and weekly.

Screenshot 1 2

Not one task. A fleet, in one shot. Joe pointed out afterward that he’d been building scheduled tasks one at a time, which is the natural instinct and also the slow way. Prompting for the whole team at once lets the model reason about how the tasks fit together instead of treating each as an isolated request.

I also told it: if you need permissions, MCPs, skills, or anything else to make these tasks succeed, tell me so we can reason through it rather than guess.

That’s the difference between an assistant and a coworker. A coworker tells you what they need.

Done is not the same as done well

AI agents keep getting smarter because the underlying models do. But their work is only as good as your ability to validate that the task was actually completed.

So every task has to be SMART. Specific, measurable, actionable, realistic, accountable, time-based. Even when the task depends on a client or another team member, there’s a way to check whether the thing actually happened.

Then the harder layer. Completion isn’t the metric. We track traffic, sales, expense, views, clicks, whatever measures whether the campaign or project actually worked. If a scheduled task doesn’t have access to the underlying measurement, it can’t tell you whether it created value or just created activity.

So I told the agent: if you don’t have access to the measurement, nudge me, follow up with me, make sure I get it to you. Otherwise we’re generating busy work with better formatting.

The agent that manages the other agents

Here’s the piece almost nobody builds.

As you work, you spawn more tasks. Those tasks overlap. The underlying skill files improve. Within a couple months you have a sprawling fleet with redundancy nobody’s auditing.

So I set up a weekly task audit. Every Monday at four in the morning, an agent reviews every task in the fleet, looks for overlap, consolidation opportunities, and refinements, and improves the whole system.

And it doesn’t just report back in the chat window. It emails me and Sean with a brutal, honest, but optimistic view of how the fleet is improving, how operations are improving, and how much work has shifted from one-off to ongoing.

I added a second one: a feedback follow-up agent. I’m a human with a lot going on, and I will miss things. If the agent flags something critical to the business, it shouldn’t say it once and let it die. It should keep surfacing it until it’s resolved.

Make page one worth reading

One standing instruction I give every agent: the first page is always an executive summary that is clear, interesting, and non-obvious. Use colors and diagrams to make the point instead of vomiting text at me.

The model can read more documents and think harder than any of us. If the output is still a wall of paragraphs, the value is trapped where nobody will dig for it.

Where all of this comes from

None of this is mine. I packaged it.

My original mentor was the CEO of American Airlines, and American Airlines pioneered Operations Research. Before the current Silicon Valley trillionaire class, the smartest people in a company sat in the OR group: the PhDs who looked at operations, engineering, manufacturing, legal, marketing, HR. The Sabre system and the credit card loyalty program both came out of that group.

What I learned there was how to run large systems that are complex, people-heavy, and full of failure modes.

Think about what can go wrong in an airline. Flight attendants strike. Fuel prices spike. Weather closes an airport. Air traffic control reroutes you. Planes need maintenance. Customers are unhappy before they board.

An airline is a giant SOP. We hope for the best and prepare for the worst. You assume things will go wrong, and when they do, the backup plan means you land thirty minutes late instead of crashing.

That’s repeatable excellence. And the same principles that keep a hundred thousand person operation running scale down cleanly to a twenty person agency. Most agencies are a good salesperson with a few big deals, an ops person, and some account managers. They never get to the point of operating like an airline. That’s the gap.

Metrics, analysis, action

The framework underneath all of it is MAA. Metrics, analysis, action.

ff

Most people stop at metrics. They pull a dashboard, feel informed, and change nothing. Some make it to analysis. Very few close the loop to action, which is the only step that changes an outcome.

It works the same way Ray Dalio’s Principles does. The whole point of that book is getting around your own blind spots. When you have a blind spot you keep making the same mistake, because your judgment is subjective and emotional. If you’re willing to look at data honestly, even when it’s ugly, you improve.

The catch is the word “willing.” Everyone claims to be logical. Almost everyone has emotional blind spots. If you’ll genuinely set your biases aside, this is physics and math. There’s nothing to argue with.

What comes next in your own system

Once the SOPs exist, the next phase is the agent coming back to you and saying: I went through everything you’ve ever done in ClickUp. These three things you do constantly are green, I can build agents for them right now. These are orange, I need a little clarification first. These I shouldn’t touch.

You say yes to the green ones. Now you’re having a conversation with team members instead of issuing commands to a search box. And nobody is running wild, because there’s a management layer QA-ing the work.

You can run this from one window. Ask it to post a detailed update in Basecamp as your assistant. It’ll ask for the connector if it needs one. You don’t have to open a single tab.

But note the qualifier. “Post an update every week in Basecamp” is lazy and it’ll produce lazy output. Actively reasoning through it with the agent is what makes it work.

Read the reasoning

Two practical habits.

Don’t scroll to the bottom. When the agent shows its thinking, read it. That’s where you learn what’s possible and what it can actually do. Skipping to the answer means you never get better at directing it.

And watch the status. A job might run three or four hours. If it looks stuck, it’s usually stuck waiting on you: a permission, a decision, a choice between two approaches. Learn the difference between working and waiting.

The chicken and egg

The reason most people never get here is that they haven’t had the first win yet. No win, no motivation, no reason to invest in the system.

So create the initial win deliberately. Take one thing you do every week, dictate three minutes of real context about it, and ask for an SOP with a measurable completion check. Then schedule it.

Once you see it come back and make you look good, the reinforcement loop starts and you’ll want more. That’s the snowball. It multiplies.

If Michael Gerber were an AI expert, this is exactly what he’d build. Functionalize everything. Take the lift off yourself.

You’re the mentor to all of these agents. The more they can see how you do things, the more they follow your example. Which is why we don’t model B players anymore.

Dennis Yu
Dennis Yu
Dennis Yu is the CEO of Local Service Spotlight, a platform that amplifies the reputations of contractors and local service businesses using the Content Factory process. He is a former search engine engineer who has spent a billion dollars on Google and Facebook ads for Nike, Quiznos, Ashley Furniture, Red Bull, State Farm, and other brands. Dennis has achieved 25% of his goal of creating a million digital marketing jobs by partnering with universities, professional organizations, and agencies. Through Local Service Spotlight, he teaches the Dollar a Day strategy and Content Factory training to help local service businesses enhance their existing local reputation and make the phone ring. Dennis coaches young adult agency owners serving plumbers, AC technicians, landscapers, roofers, electricians, and believes there should be a standard in measuring local marketing efforts, much like doctors and plumbers must be certified. He has appeared on 353 podcasts with 619 credited episodes — see the full list of his podcast appearances.