AI Adoption Is Only Part of the Value Story

AI adoption creates potential; realizing value requires changing how the business works.

That distinction has become central to my preparation for Faster Isn’t Smarter. People can use AI regularly, produce work faster, and still struggle to show what improved for the customer or the organization. Adoption is an achievement. It also leaves an important management question unanswered: what turns that activity into value?

The AWS report Reimagine: Turning AI into Value explores this gap through executive interviews and Amazon’s internal experience. One useful distinction is between measuring usage and developing the capability to work differently. Knowing that people opened a tool tells us very little about whether they challenged an assumption, redesigned a process, or produced better work. These are practitioner observations rather than proof of a universal formula, but they raise questions worth bringing into leadership reviews.

Adoption measures still have a purpose. They can reveal where access, confidence, relevance, or support is missing. The problem begins when we ask those measures to establish something they cannot: whether the business is better off.

Consider an illustrative customer-support workflow. AI helps employees draft replies more quickly. First-response time improves. More cases are marked complete.

Then follow the work downstream. Do customers contact the company again because the answer was incomplete? Are experienced employees spending more time checking recommendations? Has the queue moved from drafting to approval?

The improvement may be real. Its value depends on what happens across the complete workflow.

This is why I keep returning to a simple exercise: follow one piece of AI-assisted work to the person who receives it. Ask what became easier, what still requires judgment, and what new work appeared. A team’s productivity gain can create additional capacity—or additional work for someone else.

Ania W. Masinter’s HBR playbook, Prioritizing AI Investments That Create Real Value, makes a related point: identifying opportunities requires examining how work crosses functional boundaries. The investment decision needs to include workflow redesign, customer needs, and where competitive advantage is changing.

For leaders, that means assigning ownership beyond the tool deployment. Someone needs the authority to change the handoffs, resolve conflicting priorities, and remain accountable for the eventual outcome.

Time saved needs a destination.

Suppose a team releases several hours of capacity each week. Those hours could support additional customer demand, reduce a backlog, improve service, or make room for work that has repeatedly been postponed. They could also disappear into more meetings, more output, or fragmented gaps that are difficult to use.

The business case therefore needs a decision about what happens next. Who will redirect that capacity? What should improve? How will we know?

Multiplying estimated hours saved by salary cost may help describe potential capacity. It does not establish a cash saving. The value story needs to account for what the organization actually does with the time, alongside implementation, AI operating costs, checking, and rework.

This also connects to a distinction I have been exploring in my recent writing: AI for efficiency and AI for opportunity.

Efficiency asks how we can do existing work better or faster. Opportunity asks what becomes possible when a constraint changes.

A team that can examine customer feedback more easily might use that capability to produce its existing report faster. It might also test whether it can identify unmet needs that previously remained buried across thousands of interactions.

That second possibility requires imagination and evidence. What constraint has changed? Which customer need could we now address? What is the smallest experiment that would tell us whether the opportunity matters?

If every AI investment must justify itself through immediate labor savings, we risk narrowing the portfolio before exploring what could create growth.

The pace of learning matters as much as the pace of execution.

AI can help us make the next decision before we understand the consequences of the last one. A faster reply is visible immediately. Repeat contacts, customer trust, and the effects of a policy change may take longer to assess.

AI can help shorten some feedback loops. Other outcomes still need time to emerge.

This is where the graduated autonomy ladder in my conference work becomes useful. AI can observe, advise, prepare work for approval, act within defined bounds, or operate under delegated authority. The appropriate role depends on the decision, the consequences, and the evidence available.

Drafting a response and authorizing a refund require different permissions. Expanding authority should follow demonstrated performance, enforceable limits, and a workable route back to human review. There is no requirement for every workflow to reach maximum autonomy.

Making these choices part of everyday management matters. In Transformations That Work, Michael Mankins and Patrick Litre emphasize embedding transformation into the company’s operating rhythm, involving middle managers, and managing people’s capacity for change. Those lessons are directly relevant when employees are expected to learn new tools while continuing to deliver their existing work.

For the next AI review, I would bring together the team reporting a productivity gain, the team receiving its output, and the leader accountable for the overall result. Choose one workflow. Establish the baseline, the intended benefit, how released capacity will be used, and when enough evidence should exist to decide what happens next.

I’m particularly interested in the handoff between an AI team demonstrating a gain and a business leader making that gain useful. In your organization, who owns that handoff—and what have they needed the authority to change?

Faster Isn't Smarter: A Systems View of Decision-Making at AI Speed

An organization can make decisions faster than it can learn from them.

That is the tension behind my talk, Faster Isn't Smarter. AI can shorten the time between noticing a problem, interpreting information, choosing a response, and taking action. The consequences of that action may still take days, weeks, or months to become clear. Speed creates an advantage when we can use it well. It also makes it easier to repeat a mistake before we recognize the first one.

Imagine a customer-support team introducing AI-assisted replies. Employees respond more quickly. The queue shrinks. Managers see an encouraging early result and expand the system's role.

A week later, repeat contacts begin to rise. Some customers received incomplete answers. Others were sent to the wrong team. Experienced colleagues are now checking and correcting work that used to arrive more slowly. The first dashboard recorded a productivity gain. It did not yet capture the consequences.

This is an illustrative scenario, but it makes the management problem visible. Different parts of the same workflow run on different clocks.

The decision loop includes sensing, interpreting, deciding, acting, observing, and learning. Accelerating the first four steps does not guarantee that the last two keep pace. Sometimes AI can help there too, by surfacing patterns or connecting signals that were previously scattered. Sometimes the evidence simply takes time: a customer has to try the solution before we know whether it worked.

When leaders change direction repeatedly during that interval, they also make learning harder. Was the result caused by the new tool, a policy change, different staffing, or a change in the mix of cases? If everything moves at once, it becomes difficult to know what to repeat.

My starting point is to separate three questions that often become blurred in AI reviews.

Who is using the system? What work became faster? Did the outcome improve?

Usage tells us something about access, relevance, and adoption. Output measures tell us what people or systems produced. Outcome measures tell us whether the work achieved its purpose. Each is useful. Each answers a different question.

For the support team, that means looking beyond response time to repeat contacts, resolution errors, review effort, and total cost per resolved issue. It also means agreeing on what counts as resolved. Closing a ticket is an action in a system. Solving the customer's problem is the intended result.

One of the most useful exercises is to bring the team celebrating an AI productivity gain together with the team receiving its output. Trace an actual piece of work. Ask what the recipient needs before trusting it, where it waits, and what must be checked or redone.

That conversation can reveal a genuine improvement. It can also reveal that the bottleneck moved. Either finding gives leaders a better basis for their next decision.

The next question is what AI should be allowed to do within that workflow. I use four considerations: how quickly reliable feedback arrives, how clearly we can connect actions to results, how reversible the action is, and how serious the consequences could be.

Drafting a reply and issuing a refund are different decisions. A draft can be edited before it reaches anyone. An incorrect refund may be difficult to recover. Many small refunds can create substantial exposure before someone notices a pattern. The same model can therefore deserve different authority at different points in the process.

The graduated autonomy ladder makes that distinction practical. AI may observe, advise, prepare work for approval, act within defined bounds, or coordinate a workflow under delegated authority. Movement along the ladder should follow evidence and effective controls. Authority can also decrease when conditions change.

There is no requirement for every workflow to reach the highest level. An assistant that helps a leader challenge a strategic assumption can be extremely valuable while never being permitted to commit resources.

Before expanding authority, define what would justify the change. In the support example, compare representative cases with the existing process, agree on acceptable quality and cost, and allow enough time to observe repeat contacts. Decide in advance which errors trigger an immediate pause and who can restore human review.

This protects the learning process without requiring the organization to stand still. Teams can monitor continuously, investigate emerging issues, and intervene when harm appears. Routine changes can wait for an agreed review when doing so preserves the ability to interpret the evidence.

For leadership teams, this creates a different conversation about speed. The useful question becomes how quickly the organization can reach a trustworthy conclusion and act on it. A rapid action followed by weeks of invisible repair may be slower, in business terms, than a deliberate decision that holds.

In your next AI review, choose one celebrated speed gain and ask when its downstream consequences become visible. Does the decision to expand wait for that evidence, or has the organization already moved on?

Why you need an AI Center of Excellence

An AI Center of Excellence earns its place by connecting decisions that would otherwise happen separately.

A business team sees an opportunity. Technology evaluates a platform. Risk asks who is accountable. HR develops training. Finance asks when the investment will produce a return. These are parts of the same conversation, yet an organization can treat them as separate workstreams until something gets stuck.

I would start the design of an AI Center of Excellence with those connections. What needs to happen between an idea and a reliable capability that people can use? Where do decisions stall? Who has the authority to resolve them? The structure should make those answers clearer.

Its purpose is to help the enterprise identify valuable opportunities, deliver reliable solutions, and develop the ability to use them well. That requires attention to business value, engineering, responsible use, and how work changes. Strength in one area cannot consistently compensate for gaps in another.

A hub-and-spoke model is a useful starting point. The central team provides shared expertise, platforms, and standards. Business units own their priorities, domain knowledge, workflow changes, and outcomes. That arrangement gives teams access to capabilities they would struggle to maintain independently while keeping accountability close to the work.

At the center, I would establish four connected capabilities: AI Strategy and Value; AI Engineering and Platforms; AI Governance and Assurance; and AI Adoption and Workforce Enablement. These are responsibilities that need clear ownership. The size and maturity of the organization should determine whether they require dedicated teams or capacity from existing functions.

Strategy and Value connects the portfolio to organizational priorities. It helps leaders decide which opportunities deserve investment, identify the business owner, and define evidence of progress. Its scope should include both efficiency and opportunity: improving existing work and exploring what the business could do that was previously impractical.

Engineering and Platforms makes those choices deliverable. It connects systems and information, establishes reusable components, evaluates performance, and supports reliable operation. The work includes maintenance, cost, and recovery when something fails. A convincing demonstration leaves many of these questions unanswered.

Governance and Assurance makes acceptable use and accountability explicit. It establishes proportionate requirements for evaluation, release, monitoring, and changes in scope. Security, Privacy, Legal, Risk, and Quality bring their specialist requirements and independent challenge. Their oversight needs to remain meaningful when delivery teams are under pressure to move.

Adoption and Workforce Enablement connects the solution to everyday work. It examines which tasks and handoffs change, what employees need to understand, and how managers support the transition. Training belongs here, alongside workflow redesign and the ability to question an output or recognize when to escalate.

The reporting relationships matter. An executive sponsor gives the head of enterprise AI the authority and support to coordinate this work. An executive steering committee resolves investment priorities and competing demands. Business-unit AI leads remain accountable to business leadership while working with the CoE on shared standards and delivery. Enterprise control functions retain their oversight responsibilities.

The central team also needs boundaries. Business ownership should include responsibility for results after deployment. If a business leader wants a capability, that leader must help define success, commit operational capacity, and address the changes needed to realize the benefit. Otherwise, the CoE can become the owner of everybody else's unfinished transformation.

Consider an assistant supporting a customer-service team. The business owner identifies where employees struggle and establishes a baseline for resolution time and service quality. Strategy and Value helps test whether this is a worthwhile investment and what improvement would justify expansion.

Engineering connects approved information and builds evaluation and monitoring into the service. Governance and the relevant control functions establish requirements for access, accuracy, human review, and escalation. Enablement helps employees incorporate the assistant into their work and practice recognizing recommendations they should question.

After launch, the operating team tracks what happens. Faster drafting may leave resolution time unchanged if another handoff remains slow. Employees may save time but spend it checking unreliable recommendations. Those findings need to travel back across the same connections that supported delivery.

Leadership should make the operating agreements explicit before those situations arise. Who prioritizes investment? Who approves release? Who can accept the remaining risk, and within what limits? Who can suspend a system? Those decisions may sit with different people, and the path between them needs to be understood.

Funding needs the same clarity. Shared platforms, individual solutions, adoption, and ongoing operation all require capacity. Each solution needs owners for evaluation, incidents, improvement, and eventual retirement. A launch date should never be the point at which responsibility becomes ambiguous.

I would assess the CoE through the outcomes it helps the enterprise achieve, alongside quality, adoption, cost, and risk. Estimated hours saved need a credible path to useful capacity or financial benefit. More experiments and more users tell us something about activity; leaders still need evidence that the work is improving.

If you traced one AI initiative through your organization today, where would ownership become unclear? That is where I would begin designing the CoE.

Governance Is the Accelerator

Clear decision rights give people room to move.

I learned this while helping a regulated business respond to the spread of generative AI. People were finding uses for it across functions. The opportunity was evident, and so were questions about customer information, confidentiality, intellectual property, and responsibility. No single function owned the enterprise risk created by all those uses together.

That left us with an organizational problem. A tool could be useful to one team while creating consequences for several others. Sending each request through a succession of separate approvals would not resolve who should make the overall decision.

I brought three options to senior leadership. The decision was to widen governance so that the affected functions became the decision makers. A cross-functional task force developed the responsible AI policy, drawing on the broader company's guidance, with no incremental headcount. The CEO approved it. That policy created the governance foundation for the enterprise Copilot deployment and subsequent AI use cases.

The reframe that carried the work was straightforward: how do we enable responsible use while protecting customer information, the company, and its strategic options?

That question gave us a purpose people could work toward together. It recognized both the value of adoption and the consequences of getting it wrong. It also made clear that responsible AI use belonged in business decision-making, with the relevant expertise present from the beginning.

No incremental headcount did not mean no investment. People contributed time, knowledge, and attention alongside their existing responsibilities. The lesson I take from that experience is about using distributed expertise deliberately. It should never become an argument that governance can run indefinitely on spare capacity.

The policy was a foundation. Turning it into a consistent operating process remained work in its own right. My later governance materials addressed intake, prioritization, ownership, production support, and lifecycle decisions. That distinction matters: leadership approval of a policy and reliable execution of that policy are separate achievements.

For leaders building their own model, I would start with the decisions the governance body needs to make. Which risks require collective judgment? Which investments require enterprise prioritization? Which matters can be handled by an accountable owner within agreed boundaries?

A council needs a clear remit and the authority to resolve those questions. Membership should follow the decisions and their consequences. Business, technology, security, privacy, legal, quality, finance, and procurement may all be relevant, depending on the use case. Attendance alone does not create shared accountability. People need to know what they are there to decide.

Operational review also needs a home. Someone must help requestors describe the problem, identify missing information, bring in the right expertise, and maintain the decision record. Senior leaders should receive the unresolved tradeoffs that require their authority, with enough context to make a decision.

This is where intake can either help or obstruct. A useful intake asks what outcome should improve, who owns it, what data is involved, what the system will be permitted to do, and who could be affected by an error. It also checks whether an existing capability already meets the need.

A familiar tool used with public information and a familiar tool used to influence a sensitive decision may require very different treatment. Approving the product name alone leaves too much unresolved. The use, the information, and the authority all matter.

The path through governance should be proportionate. Teams working within an established pattern should understand the conditions for proceeding. Unusual data, higher consequences, new external commitments, or expanded authority should trigger additional review. Clear routes reduce the need to renegotiate the same questions every time.

Decisions need to travel with the work. Record what was approved, for which purpose, under what conditions, and who owns the next step. A requestor should be able to understand whether the answer is proceed, revise, test further, use an existing solution, or stop. Ambiguity creates waiting and encourages workarounds.

Then keep governance connected to the system after launch. A change in data, model, vendor, audience, or permissions can change the basis on which the original decision was made. Ownership, monitoring, incident response, and retirement therefore belong in the operating model from the start.

The same applies to business value. An approved system that no longer serves a useful purpose still consumes money and attention. A governance process should make it possible to change or retire that system, as well as authorize it.

I would assess governance by the quality and clarity of the decisions it enables. Do teams know where to begin? Are familiar requests handled consistently? Do exceptions reach someone with authority? Can the organization explain why a system is allowed to operate and recognize when those conditions no longer hold?

Those are practical tests of whether governance is helping people act responsibly. In your organization, which decisions keep returning for approval because the underlying authority was never made clear?

In the Era of AI - Judgment Is Becoming the Scarce Resource

AI is making judgment more valuable as it loosens the constraints that used to force decisions.

Leaders have always worked with incomplete information. Time, analytical capacity, and access to information placed practical limits on how much we could explore. Eventually, gathering more became too expensive or too slow. We had to choose what mattered, accept some uncertainty, and act.

Those constraints did some of our prioritization for us. A team could only investigate so many opportunities, develop so many scenarios, or prepare so many recommendations. The limits were often frustrating, but they gave the work a stopping point.

AI changes some of that. We can ask another question, generate another option, challenge another assumption, or request another version at a much lower cost. Time and resources still matter. But in parts of knowledge work, the ability to produce is becoming less of a constraint than our ability to decide what deserves attention.

I see three connected leadership questions emerging from that shift: what is worth doing, when do we know enough to act, and how do people develop the judgment to make those calls?

The first is about how we allocate the capacity AI creates. When something becomes easier and cheaper to produce, we can simply produce more of it. A presentation takes an hour instead of a day, so we create more presentations. An analysis becomes easier to run, so we explore more scenarios. Each additional piece may be useful. Collectively, they can expand the work without improving the outcome.

The cost of generating something also tells us little about the attention it will consume. Someone still needs to read the presentation, test the assumptions, reconcile conflicting recommendations, and decide what happens next. A team can become faster at producing work while making the surrounding organization busier.

That makes the allocation decision more deliberate. Capacity could go toward serving customers better, investigating a neglected opportunity, improving quality, or giving people room to learn. It can also disappear into additional output. Leaders need to decide what the freed capacity is for and what evidence would show it was used well.

The second question follows closely: when is more analysis no longer useful enough to justify delaying a decision?

We have spent a long time learning to decide with less information than we would like. Increasingly, we may also need to recognize when we have enough. There will almost always be another scenario to model or another assumption to examine. The availability of more analysis does not tell us whether it could change the choice.

Consider a team comparing two ways to improve a service. It has a clear objective, a credible baseline, and a small trial it can reverse. Another round of analysis might refine the forecast. Running the trial might resolve the uncertainty that matters. The judgment lies in recognizing which next step will teach the team something useful.

I would make that discussion explicit before asking for more work. What decision are we making? Which uncertainty could change it? What evidence would be sufficient to proceed? How costly would it be to discover we were wrong? A reversible experiment and a consequential commitment deserve different thresholds.

AI can help examine these questions, but someone still needs to own the stopping point. Otherwise, analysis can become a comfortable way to postpone responsibility. The recommendation gets more polished while the underlying decision remains untouched.

The third question is the one I find most challenging. Judgment becomes more valuable at the same time that AI may be changing how we learn to exercise it.

Much of professional development has happened through doing the work. Investigating a problem, building an analysis, making a recommendation, and seeing what happened helped people develop context and pattern recognition. Some of that work was repetitive. Some of it also exposed the details that made a later judgment possible.

If AI takes on more of the production, we need to be deliberate about what replaces those learning opportunities. Asking someone to evaluate an answer assumes they have a basis for recognizing what is missing, implausible, or misleading. A polished output can make that harder, especially when the person reviewing it has limited experience with the underlying problem.

There are practical ways to preserve learning while changing the work. Ask people to form an initial view before consulting AI, explain which evidence they trust, and compare competing recommendations. Let them make bounded decisions, observe the consequences, and discuss what they missed with someone more experienced. Increase responsibility as their judgment becomes more reliable.

That requires time from managers and experienced colleagues. Organizations should account for it when deciding how to use AI-created capacity. If every available hour goes into additional output, there may be little room left to develop the people expected to evaluate it.

These three questions belong in the same leadership conversation. We need judgment to choose the work, to decide when the evidence is sufficient, and to help others learn to do both. As production gets easier, I am increasingly interested in where organizations are deliberately creating those opportunities to exercise judgment—and where they are quietly removing them.

Email is becoming less of a communication tool and more of an organizational memory system.

Email is becoming less of a communication tool and more of an organizational memory system.

TL;DR: AI agents are increasingly reading, summarizing, retrieving, and acting on email before humans do. That changes email’s function, its audience, and eventually the role communication plays at work.

Tom Fishburne’s “AI Written, AI Read” cartoon captures the emerging absurdity perfectly: one AI expands a simple thought into an email. At the same time, another AI condenses it back into the original thought.

For nearly fifty years, email has served as a digital replacement for letters and memos. We write messages, send them to people, and expect those people to read, interpret, and respond. But the primary reader is beginning to change. Many people, including me, now use AI to summarize inboxes, prioritize messages, draft responses, extract commitments, schedule meetings, and surface action items.
The email itself has not changed. Its function and audience have.
Email is becoming a record of organizational decisions, approvals, obligations, and context. AI can now interpret that record at scale. Instead of asking a colleague to find and forward an email from six months ago, we can ask an agent:
“Summarize every discussion we have had about this idea, identify the commitments we made, and tell me what remains unresolved.”
That shift means we are increasingly writing for two audiences: the human recipient and the AI systems that may later retrieve, categorize, summarize, and reason over the message.
Good workplace writing may therefore become more structurally explicit: clear decisions, named owners, specific deadlines, linked context, visible assumptions.
The inbox may disappear before email does.
A significant share of workplace email exists primarily to coordinate people and activities:
“Any updates?”
“Following up.”
“Just checking.”
“Circling back.”
“A friendly reminder…”
As agents become better at tracking shared work, some of this coordination may no longer require a message at all.
What remains human is not simply the transmission of facts. It is negotiation, coaching, conflict resolution, persuasion, trust building, and the creation of shared meaning.
The broader question is not whether AI will help us write more email. It is what human work becomes when we spend less time sending, receiving, forwarding, summarizing, and searching for information.
Our roles may shift from carrying information through the organization to interpreting it, exercising judgment, and helping people decide what to do next.
#FutureOfWork #ArtificialIntelligence #OrganizationalDesign #KnowledgeManagement #AgenticAI

Building an agent makes leadership assumptions turn into insights

I spent eight weeks building the agent system that now supports my own work. That experience changed the questions I bring to AI. Words such as context, permissions, memory, and oversight became choices I had to make in order for something useful to work reliably.

It is easy to agree that an agent needs good context. Building one forces you to decide which information it should use, where that information lives, how current it needs to be, and what happens when two sources disagree. The concept becomes a practical responsibility.

The same happens with authority. Asking a system to help manage your day sounds simple until you separate reading your calendar, preparing a plan, creating a task, changing an appointment, and contacting another person. Each step has a different effect. Each deserves a deliberate decision about permission.

That is why I encourage leaders to build a bounded prototype themselves. A small build gives them something concrete to inspect, question, and improve. They can experience the distance between a plausible demonstration and a workflow they are willing to depend on.

My morning-brief playbook contains an example of that distance. It requires the system to name sources it could not read. A missing source must remain visible in the brief. Otherwise, an empty section can look like evidence that nothing needs attention when the system simply failed to check.

That is a design decision about trust. It is also an organizational lesson. In a business workflow, a reassuring summary can conceal an information gap unless someone has specified how absence and failure should be represented.

A useful leadership workshop can begin with something equally bounded: preparing a meeting brief from a small set of approved documents. The initial goal is specific. The agent may read those documents and create a draft. It cannot contact attendees, change a system of record, or make a commitment on the team's behalf.

Participants then define quality. What must the brief include? Which statements need a source? How should it handle conflicting dates or an unanswered question? What would make the result unusable, even if the writing sounds excellent?

Next, test the boundaries. Remove an important document. Insert an outdated version alongside a newer one. Include a request in the source material that conflicts with the assigned task. See whether the system identifies the problem, invents a resolution, or proceeds beyond its authority.

This is where the learning becomes memorable. The leader has to explain what the system should have done and translate that expectation into the next version. A vague preference for accuracy becomes a concrete rule about sources, uncertainty, or escalation.

The exercise also reveals the components of the system. There is a purpose and role. There is context. There are repeatable instructions, tools, permissions, and a place for outputs. There may be retained information between runs. There must be a way to evaluate performance and respond when something changes.

Understanding those components makes conversations with technical teams more productive. A leader can ask whether the problem comes from missing information, an unclear workflow, an unsuitable tool, or an undefined acceptance criterion. Asking for a better model is only one possible response.

The prototype will also expose what belongs in the surrounding organization. Who owns the source documents? Who can resolve a policy conflict? Who maintains the workflow? Who receives the exception? An agent cannot make these ownership questions disappear simply by presenting a confident answer.

There is a limit to what the exercise establishes. A prototype that works for one person with a few documents has not demonstrated readiness for enterprise use. Broader access, multiple users, sensitive information, integration failures, and operational support introduce further requirements. The value of the build is that leaders can now see and discuss those requirements more clearly.

The other change is more expansive. Once I can make a small system, I can ask different questions about what is possible. I can move from accelerating a familiar task to experimenting with a different way of achieving the outcome.

A workshop need not end when the presentation ends. Participants might revisit a scenario, test a decision, and receive feedback on their reasoning. A recurring report might become a conversation about exceptions and choices. These are possibilities to explore, with their own evidence requirements, rather than guaranteed benefits of adding AI.

That is the connection between building and imagination. The act of making something tests both the technology and the assumptions that shaped the work. It creates an opportunity to ask which constraints are real, which have changed, and which we have stopped questioning.

For your next leadership discussion, bring a small working example and the cases where it failed. Which conversation becomes possible when everyone can inspect the same behavior—and which idea might you now consider that a slide presentation would never have surfaced?

Women Building with AI: Participation Shapes the Decisions

If you asked me who is the most productive and efficient person I know, I would pick my mom, and then my sister. OK, I know that’s two people, and then I would have to add everyone else who is juggling their personal and professional lives and making it all look easy. Which is why I’m so thrilled to see more and more women building with AI.

Many women already know exactly what they would change about the way work gets done.

They know which process requires too much chasing, which information is always missing, and which task takes an hour when it should take ten minutes. They have developed workarounds, connected the pieces, and kept things moving. That knowledge is a valuable starting point for building with AI.

What may be less familiar is the possibility of building something themselves.

I spent eight weeks building the agent system that now supports my own work. I started with problems I understood: work I repeated, information I needed, and coordination that consumed attention I wanted to use elsewhere. Building helped me turn those observations into something I could test and improve.

I could examine which information the system needed, which actions I was willing to delegate, and what evidence would let me trust a result. When something did not work, there was a specific behavior to investigate. The conversation became more precise because I had something real in front of me.

For women whose expertise is in operations, strategy, finance, people, customer experience, or a particular industry, that expertise is a starting asset. Understanding the work helps us notice when an output misses the point, when an exception matters, and when a proposed efficiency creates difficulty for someone else.

Those judgments shape a system. So do choices about what success means and who gets to challenge the result. Building creates a way to express professional knowledge in the design, before a process hardens around someone else's assumptions.

You do not need to begin with an ambitious product idea. Start with something you regularly find yourself thinking: “There has to be a better way to do this.” Describe what happens today, what you wish happened instead, and what would make the result useful enough to trust.

Choose a first experiment you can test without exposing sensitive information or affecting another person. A meeting-preparation aid using public material, a planning exercise with sample data, or a tool for organizing your own learning can provide a useful place to begin.

Decide what it may read, produce, or change. Then try it and inspect what happens. You do not need to understand every technical component before starting a bounded experiment, but you do need to understand the consequences of the permissions you grant.

Invite another person to test it. Ask them to use incomplete information, challenge an answer, or try a case you did not anticipate. Their experience can reveal what you assumed was obvious. Learning how to improve the design is part of the accomplishment.

This also changes how we talk about confidence. A useful basis for confidence is evidence: I built this, tested these conditions, found these limits, and know where I need help. That position leaves room for curiosity and collaboration. It does not require pretending to have mastered a field that is still evolving.

The opportunity extends beyond personal productivity. A faster research brief may be helpful. A new way to prepare for a difficult decision, teach a concept, explore a customer need, or test a business idea may change what you are able to attempt.

I want more women involved in that exploration. The question is larger than whether AI can help us complete today's task list. It includes which possibilities we recognize, whose needs we consider, and which changes we decide are worth pursuing.

Responsibility for participation also sits with organizations. Encouraging women to experiment is insufficient if access, protected time, sponsorship, and meaningful ownership remain uneven. Leaders can create paid opportunities to build, provide appropriate tools and support, and make sure the people doing the work receive credit for it.

That includes participation in consequential decisions. Who defines the evaluation criteria? Who presents the work to leadership? Who owns the pilot and receives the next opportunity? Being invited to test a tool is different from having a voice in how the organization uses it.

Peer communities can help people enter and remain in that work. Small groups can compare what they are building, make uncertainty discussable, and help each other resolve practical obstacles. The most useful support may be a colleague willing to sit beside someone while they try something unfamiliar.

There should also be room to decline a particular tool or use case for a sound reason. Participation does not require accepting every vendor claim or granting every requested permission. The ability to examine an approach and explain why it should change is itself a form of influence.

For women leading teams, I would make one small build a shared learning experience. Choose a real question, create a bounded experiment, and discuss what it reveals about the work. For leaders sponsoring that effort, make the time and decision-making access real.

Your experience gives you a place to start. Building lets you discover what else is possible.

If you aren’t already building, reach out if you'd like help getting started!

The Work of the Future: Deciding What Deserves to Be Done

As production gets easier, choosing the work becomes more consequential.

A team can use AI to produce more analyses, proposals, summaries, and plans. Each item can look useful in isolation. Together, they can create a growing demand on everyone else's attention. Someone still has to read, judge, connect, and act on the work.

This is the part of the future-of-work conversation I find most interesting. Expanding what a team can produce also expands the set of things it can choose to do. Leadership has to make that choice more deliberately.

Consider an illustrative weekly reporting process. AI reduces preparation time dramatically. The natural response is to make the report more detailed or produce it more often. Before doing either, I would ask who uses it and which decision it changes.

If the report primarily reassures people that work is happening, faster production may preserve a weak process. If it helps someone allocate scarce capacity, the opportunity might be to surface the few changes that require a decision. The useful redesign depends on the purpose.

This calls for looking at the portfolio of work, including the small recurring obligations that rarely appear in a strategic plan. Meetings, status updates, reconciliations, internal presentations, and requests for information all consume attention. Making them cheaper to produce can increase the volume unless someone changes the demand.

There is a practical question worth asking before automation: if we stopped this activity, what would become worse, for whom, and how soon would we know?

Some answers will reveal essential work whose value has been underappreciated. Others will expose a report without a reader, an approval without a decision, or a meeting that exists because information is difficult to find. AI can help us investigate those patterns, but the choice about what deserves to continue remains a leadership responsibility.

I also want to protect space for a more ambitious question: what have we not attempted because the effort seemed disproportionate to the opportunity?

In my recent writing, I have been exploring AI for efficiency and AI for opportunity. The first lens examines existing work. The second examines possibilities that a changed constraint brings within reach. Both belong in the conversation.

Imagine a professional community with years of useful discussions that are difficult to navigate. An efficiency approach might summarize each new session more quickly. An opportunity approach might test whether members can explore the accumulated knowledge around their own challenges, with clear sources and a route to human discussion.

The second idea introduces work the organization was not previously doing. Its value would need to be tested: do people find relevant knowledge, apply it, and make better decisions? The point is to give that possibility a fair hearing before every available hour is absorbed by existing routines.

That makes released capacity a strategic resource. A leader can explicitly allocate some of it to service improvement, learning, or experiments. Without that choice, the organization may simply add more assignments to the same people and call the result transformation.

There is also a question about professional development. Much expertise grows through doing work, receiving feedback, and recognizing patterns over time. If AI takes over parts of that work, leaders should examine how people will continue to develop the judgment needed to supervise it.

For example, a junior colleague who always receives a polished analysis may have fewer opportunities to notice how an ambiguous question was framed. A redesigned learning path could ask them to make an initial assessment, compare it with the AI-assisted version, and explain the differences. The work changes, and the learning design should change with it.

Performance expectations need attention too. If employees are rewarded primarily for volume, making production easier can encourage more output regardless of its usefulness. Measures of customer impact, decision quality, resolved problems, and reduced downstream effort give a different signal about what the organization values.

This does not mean every activity needs an immediate financial return. Exploration, relationship-building, and professional learning can be worthwhile before their effects are measurable. Leaders should describe why that work matters and what evidence would justify continuing, changing, or ending it.

A practical starting point is a work review alongside the technology review. Select one team and examine what should continue, what should change, what can stop, and what newly possible work deserves a test. Invite the people receiving the team's output, because they can often explain its value more clearly than the people producing it.

I would expect that conversation to be uncomfortable in useful ways. It may expose an executive request that no longer serves a purpose. It may reveal that the most valuable use of AI is to make space for work leadership has repeatedly postponed. It may show that a slower, more thoughtful exchange deserves to remain human.

If your team could produce twice as much next quarter, which work would you deliberately decline to expand—and what would you finally make room to try?

Measuring AI Value: What the Usage Dashboard Can—and Cannot—Tell You

An AI dashboard can show active users, completed training, prompts submitted, and hours reportedly saved. Those numbers can help a program team understand reach and engagement. They do not, by themselves, establish that a customer received better service or that the business realized a return.

My concern is the moment when evidence changes meaning as it travels upward. A user reports finishing a task faster. A project team estimates recovered capacity. A leadership presentation labels it savings. Each step sounds plausible, but the final claim may require decisions and evidence that never existed.

I would begin with the intended outcome and work backward. Who should be better off because of this investment? What should change for them? Through which changes in work would that benefit occur? This creates a value hypothesis that can be tested.

Consider an illustrative service workflow. The hypothesis might be that AI helps employees prepare accurate responses more quickly, reducing total effort per successfully resolved issue while maintaining service quality. Usage, speed, quality, and cost now have distinct roles in evaluating that hypothesis.

Usage tells us whether enough eligible work reaches the system for the proposed benefit to be plausible. Output measures show whether responses are produced more quickly. Quality measures show whether those responses resolve the issue. Cost measures include the effort and resources needed across the full workflow.

If employees use the system but outcomes do not improve, that is useful evidence. The tool may be helping with the wrong step, creating additional review, or producing work that downstream teams cannot use. If outcomes improve but adoption is narrow, the next question may concern access or applicability rather than model quality.

Leading indicators help leaders act before the final outcome is available. For example, the share of eligible cases with complete information may tell us whether a workflow is ready to perform. The rate at which users abandon a draft may reveal a problem that deserves investigation. Neither measure should be promoted into proof of business impact without checking the relationship.

A leading indicator is a hypothesis about what precedes the outcome. Its usefulness should be revisited as evidence accumulates. A measure that once identified a barrier may lose relevance after the process changes.

Balance measures also matter. Faster responses can be paired with repeat contacts and errors. More content can be paired with evidence of usefulness to its audience. Reduced processing time can be paired with review and correction effort. The aim is to detect whether an apparent improvement shifts cost or difficulty elsewhere.

Time saved needs especially careful treatment. Estimated hours can indicate capacity. Realized cost reduction requires an actual change in spending. Growth requires evidence of additional business. Better service requires an appropriate customer or operational measure. These benefits may be related, but they should not be counted as interchangeable.

Suppose a hypothetical team releases ten hours a week. If those hours are used to clear a backlog, track the backlog and the quality of completed work. If they allow the team to absorb more demand without additional staffing, describe that as capacity or avoided future cost, with the assumptions visible. If they support experimentation, explain what was learned and what decision it informed.

The same ten hours should not quietly appear as a cash saving, extra revenue, and increased capacity in three different reports. A benefit owner, working with finance where appropriate, should reconcile the claims and identify which benefits have actually been realized.

Costs need an equally complete view. Include implementation, integration, data preparation, licenses or model use, support, review, rework, and maintenance. Some costs will be shared across several use cases. Make the allocation clear enough that comparing projects does not reward whichever team can move expenses out of its own budget.

Then ask what caused the change. Compare similar work and account for differences in case mix, staffing, seasonality, or simultaneous process improvements. A controlled test is useful when practical. A staged rollout or careful before-and-after comparison may be more feasible, with the limitations stated honestly.

Choose an observation window that matches the outcome. A response can be measured immediately. Repeat demand and lasting behavior change take longer. Early evidence can justify continuing an experiment without justifying a full return-on-investment claim.

This is also why I would preserve evidence about the decision itself: what was expected, which assumptions mattered, and what would change the investment. A disappointing result can produce valuable learning. A favorable result can occur for reasons the team did not anticipate. Both deserve examination.

For your next value review, choose one number in the headline and trace it back to the original observation. Which parts are measured, which are estimated, and which depend on a management decision that still has to happen?

From AI Literacy to AI Mastery for Knowledge Workers

An introductory session can create awareness. A demonstration can create enthusiasm. The harder question is what someone can do the following week, when their information is incomplete, the example does not fit, and the person delivering the training is unavailable.

My work on an agent-building cohort made that question concrete. In the recommendations and diagnostic I developed, I looked beyond the content to the conditions around learning: setup, confidence, relevance, time, practice, feedback, and support after the course ended.

The useful question was whether the participant could keep progressing. A person who needs help installing a tool has a different problem from someone who cannot identify a worthwhile use case. Someone concerned about permissions needs a clear explanation and an appropriate way to experiment. Another video will not resolve all three situations.

Enterprise AI literacy has the same challenge. A single training pathway can hide important differences between roles, starting points, and responsibilities. People need a common foundation, followed by opportunities to practice the decisions their work requires.

For employees using AI to prepare content, that may include selecting appropriate information, checking claims, and recognizing when an output is unsuitable. For managers, it includes redesigning work, assessing results, and making space for learning. For builders and owners, it includes permissions, testing, exceptions, maintenance, and the consequences of changes.

Leadership literacy also deserves attention. Senior leaders need enough direct experience to question an impressive demonstration, understand what a usage metric establishes, and recognize when a proposed benefit depends on an organizational change. Delegating technical implementation does not remove the need to understand the decisions being made.

I would measure progress through demonstrated capability. Can the person define a useful task, provide appropriate context, evaluate the result, and explain what still requires judgment? Can they recognize a boundary and ask for help? These are more informative than knowing whether they attended a session.

That changes the design of learning. Begin with a worked example, give people a partially supported task, and then ask them to adapt the approach to their own work. Include a result that needs correction. The moment someone explains why an output is inadequate can reveal more understanding than a polished demonstration.

Champion networks can help this practice travel across the organization. Their value comes from proximity to real work. A colleague who understands the team's deadlines, information, and recurring frustrations can translate a general capability into something relevant.

Choose champions for curiosity, sound judgment, and willingness to help others, alongside technical confidence. Give them time and access to support. If the role exists only as an additional expectation on already busy people, the organization is relying on personal generosity to sustain an enterprise capability.

Be explicit about the role's limits. Champions can help colleagues discover uses, practice, and surface problems. They should have a route to people who can resolve questions about access, policy, quality, or integration. They should not become an unofficial approval authority simply because everyone knows to ask them first.

A community of practice gives those experiences somewhere to accumulate. Its most useful sessions can center on an actual piece of work: what someone tried, what happened, where the output failed, and what they changed. A failed experiment with a clear lesson can be more valuable to peers than another flawless demonstration.

Capture reusable knowledge carefully. A shared example should explain its purpose, inputs, limits, and how to judge the result. A collection of prompts without that context can be difficult to adapt responsibly. Give someone responsibility for maintaining examples as tools and policies change.

There is an opportunity to use the network as a source of organizational intelligence. Repeated questions may reveal an unclear policy, inaccessible information, or a process that needs redesign. Those patterns should reach the teams with authority to change the conditions, rather than leave champions answering the same question indefinitely.

Managers make reinforcement credible. They can allocate time for practice, discuss how work changed, recognize thoughtful evaluation, and make it safe to report that an AI approach was unhelpful. If every conversation celebrates usage, people may learn to conceal the evidence that would improve the program.

I would also check what remains in use after the initial support ends. In my course diagnostic work, I proposed a later pulse asking whether participants were still using what they had built and what had broken or been abandoned. The answers would guide support and reveal whether the program had changed behavior.

At enterprise scale, that follow-through belongs in the operating rhythm: review useful practices, identify recurring friction, refresh guidance, and adjust learning to what people are actually trying to accomplish.

Think about the last AI training your team attended. What can they now do independently, and where would they turn when the example stops matching the work? Those two answers tell us a great deal about what the literacy program needs next.

Agentic AI Operating Models: Graduated Autonomy Ladder

An agent's completion message is a claim that the operating model must be able to check. So how can we evaluate how well your agents are working?

Building my own agent system made this concrete. My monitoring page distinguishes a recorded run from a missing run and from a workflow for which no expectation has been configured. It also shows when the checker last ran. A reassuring status loses meaning if the mechanism producing it has stopped.

Those distinctions are deliberately modest. Evidence that a routine ran does not prove that its output was useful or complete. A delivery record does not establish sound judgment. Each additional claim needs its own evidence.

This is a useful starting point for enterprise agentic operating models. Before expanding what agents can do, establish what the organization can observe about their work and who will act on that information.

An operating model should describe how a particular outcome is produced. Where does the request originate? What work is assigned to people, fixed automation, or agents? What information and systems are involved? Who receives the output? Who is responsible when the chain stops halfway through?

The model needs to remain understandable beyond the team that built it. A business owner should be able to describe the service and its limits. An operator should be able to identify a failure and take the next step. A reviewer should be able to inspect the evidence relevant to the decision they are being asked to make.

Take an illustrative intake agent. It reads a new request, checks required information, searches an approved catalog, and prepares a recommendation for routing. A second component updates the record after the authorized person makes a decision. The workflow crosses several boundaries, even though the user experiences it as one service.

At each boundary, something should be verifiable. Was the correct request read? Which catalog version was available? Was an update accepted by the receiving system? Did the recommendation reach the assigned reviewer? If a step failed, did the workflow stop or continue with incomplete information?

The relevant record should contain observable actions, source references, applied rules, approvals, results, and exceptions. A generated explanation can help a reviewer understand a recommendation, but it should not be mistaken for proof of what the system actually did.

Authority then becomes specific. The agent may read certain records, create a draft in a designated location, and route eligible requests. It may be prohibited from approving investments, changing policy, or contacting people outside the workflow. These boundaries need to be reflected in permissions and system behavior, as well as instructions.

A named business owner remains accountable for the outcome. That person does not have to perform every technical task. Technical operations may maintain integrations and reliability; data owners may maintain sources; control functions may define mandatory constraints. The operating model should make those responsibilities connect rather than leave gaps between them.

Exceptions deserve the same care as the expected path. Define who receives an unresolved case, what information accompanies it, how urgently it needs attention, and what happens if the recipient is unavailable. A handoff is incomplete until someone can accept responsibility for the next action.

Recovery requires choices too. If the system attempted a write and received no confirmation, should it retry? Could that create a duplicate? If only part of a workflow completed, can the team safely resume it? These questions are easier to answer before an incident, with a clear record of what is known and what remains uncertain.

NIST's AI Risk Management Framework treats governance, measurement, and management as continuing responsibilities across the lifecycle. That is a useful orientation for agents: authorization begins an ongoing responsibility to assess performance and respond to change. The framework is voluntary guidance, and its application needs to fit the organization's context. Source

Widening authority should therefore be a documented decision about a defined class of actions. State which evidence supports it, which limits remain, and what would trigger suspension. An agent that performs well on one class of request has not automatically demonstrated readiness for another.

Changes to the system can also invalidate prior confidence. A new model, connector, source document, or business rule may alter behavior. Retest relevant cases, retain a record of the change, and decide whether the current permissions still make sense. The organization should be able to reduce authority as readily as it expands it.

The commercial side belongs here as well. Someone needs visibility into operating cost, human review effort, incident handling, and continued usefulness. A service that costs more to maintain than the value it creates should be eligible for redesign or retirement, even if its technology remains impressive.

Before approving the next expansion of an agent's role, ask the team to reconstruct one completed case and one failed case from the available evidence. What can they establish, what must they infer, and which missing observation would change your willingness to delegate?

An AI transformation becomes durable when the new way of working can survive an ordinary week.

Launch week is a forgiving environment. The project team is available. Sponsors are attentive. Early adopters are motivated, and unusual problems receive immediate help. The more revealing test comes later, when someone is absent, priorities compete, and the person who designed the workflow is no longer sitting beside the user.

That is the test I would design for from the beginning. Can people complete the work? Do they know what to do when the system is wrong? Does the new process still make sense when the organization is busy?

My enterprise work and more recent learning-program design keep bringing me back to sequencing. Teams need enough clarity, capability, and support to take the next step. A roadmap that asks them to change everything simultaneously can make even a useful technology difficult to absorb.

Start with a business outcome and a bounded workflow. Choose something meaningful enough to attract ownership and specific enough to understand. Map who does the work, where information comes from, which decisions require judgment, and where people wait or redo someone else's output.

Include the person who receives the result. An internal team may describe a task as complete at the point another team's effort begins. That handoff is often where the transformation needs attention.

Then establish the conditions for an experiment. Identify the owner, baseline, available data, permitted uses, and support arrangement. Decide what success would look like and what could make the test unacceptable. Some gaps need to be fixed before the experiment; others can be investigated through it. The important distinction is which uncertainty the organization is deliberately testing.

A useful early choice combines business relevance with a manageable consequence of error and a reasonable path to feedback. That does not mean every transformation must begin with a trivial use case. It means the organization should understand what it is learning and be able to respond when the evidence challenges the plan.

Redesign the work alongside the technology. If AI prepares a first draft, who checks it? If information arrives more quickly, does the decision still wait for a monthly meeting? If the new workflow removes a coordination step, which role now owns the exception that step used to catch?

These questions can lead to changes in approval rights, team responsibilities, service expectations, and performance measures. They are operating-model decisions. Leaving them unresolved can make the AI system appear to fail when the surrounding process has never been adapted to use it.

Build adoption support around real barriers. In my draft capability diagnostic for an agent-building cohort, I separated concerns about technical setup, permissions, time, and the absence of a useful personal use case. Each calls for a different response. More training content will not create time in someone's calendar or resolve a concern about connecting sensitive information.

The same principle applies at enterprise scale. An employee may understand the tool and still have no reason to change a familiar method. Another may see the value but need practice. A manager may support the initiative while continuing to reward the old behavior. Treat these as different design problems.

During the pilot, study the work people actually do. Which steps do they skip? Where do they copy information manually? Which outputs do they distrust? Which unofficial workarounds keep the process functioning? These observations can reveal requirements that a successful demonstration never exposed.

Plan the transition away from the old process. Running both indefinitely creates duplicate effort and ambiguity about which record is authoritative. There may be a good reason for a temporary parallel period, but it needs an owner, a purpose, and criteria for ending it. Otherwise, employees experience the transformation as an additional obligation.

Expansion should follow readiness across the workflow. Does the next team have the same data quality, skills, permissions, and support? Will greater volume overwhelm reviewers? Can the system handle the exceptions that were rare in the pilot? Scaling changes the conditions, so the evidence needs to travel with the rollout.

There is also a portfolio question. The same employees may be absorbing several changes at once. Leaders should make the demand on their attention visible, sequence competing initiatives, and identify work that can stop. A transformation plan that assumes unlimited organizational capacity is making a hidden resource decision.

Finally, give reinforcement an owner. Maintain the guidance, keep peer support active, review failures, and revisit whether the benefit is appearing. In the course recommendations I developed, the question I wanted to ask after the program was whether the participant's agent was still being used and what had caused them to stop. Enterprise adoption deserves the same curiosity.

For a workflow you have already launched, look at the first month after the extra project support ended. What continued to work, what quietly reverted, and what does that tell you about the next change the organization is ready to absorb?

Assistant or Agent: What Changes When AI Starts Doing the Work

Delegating a task to AI changes the work of the person responsible for it.

An assistant can help me prepare a meeting brief. I bring the question, review the response, ask for revisions, and decide how to use it. An agentic workflow can take on more of the sequence: find relevant material, check several sources, assemble the brief, place it where I need it, and report what it could not complete.

That sounds like a difference in convenience. For a leader, it is also a difference in responsibility. Someone now has to specify what counts as completion, which sources are appropriate, what may be changed, and what should happen when the expected path breaks.

The names are not reliable enough to answer those questions. Products described as assistants may use tools and perform actions. Systems called agents may follow tightly prescribed workflows. I care more about the delegation: what can this system do without a person directing every step?

Anthropic's Building Effective Agents makes a useful architectural distinction between predefined workflows and agents that dynamically direct their processes and tool use. Its advice to begin with the simplest adequate approach is relevant here. A fixed sequence of known rules may be better served by conventional automation. More autonomy needs a reason. Source

My practical test starts with the work itself. Does it mainly require an answer, interpretation, or creative exchange that a person will use? An assistant may be sufficient. Does it require pursuing a defined outcome through several steps, responding to what happens along the way? An agentic approach may be worth examining.

Then comes the harder question: can we define a boundary within which that pursuit is useful and acceptable?

Consider a proposed intake workflow. A request arrives with a missing business owner and an ambiguous description. The system could identify the missing field, compare the request with the existing catalog, and prepare a clarification for the requestor. Those are distinct responsibilities. Approving the underlying investment is another responsibility entirely.

In a course design exercise, I explored this kind of separation for AI intake and triage. The useful design question was where routine coordination could be reduced while preserving human judgment for substantive risk and investment decisions. It was a proposed design, not evidence that a deployed agent had achieved the hoped-for result.

The exercise also makes a broader point: a process rarely belongs wholly to either a person or an agent. Responsibility can differ from step to step. A system may retrieve information independently, propose a classification, prepare a record for review, and stop before making a commitment.

Before it acts, it needs explicit permissions. Reading a record, editing a record, sending a message, and approving a request are separate capabilities. Access should be limited to the task, and consequential boundaries should be enforced through the surrounding system. A sentence asking an agent to be careful cannot carry the entire control burden.

It also needs a usable definition of quality. A meeting brief should identify its sources, separate confirmed facts from inference, and expose missing information. An intake assessment should apply the right criteria and route uncertain cases appropriately. Fluency does not establish that either job was done well.

Testing should include incomplete inputs, conflicting records, unavailable tools, duplicate requests, and plausible attempts to redirect the system outside its purpose. The question is how it behaves when conditions depart from the demonstration. A useful agent should sometimes stop, ask, or return a partial result with its limits made clear.

That creates work for people. Someone has to receive exceptions, understand them, and resolve them within the time the workflow requires. If every unusual case accumulates in an unattended queue, the organization has automated the easy steps and left the process incomplete.

Human review needs design as well. Reviewers need the evidence, context, time, and authority to challenge the recommendation. A person clicking approve under pressure may add very little protection. For important decisions, it can be useful to compare an independent human assessment with the AI recommendation before assuming agreement demonstrates quality.

Completion must also be visible. Producing a draft does not establish that it was saved. Preparing an update does not establish that the receiving system accepted it. A reliable workflow distinguishes attempted actions from confirmed results and reports where the chain ended.

Finally, delegation should change deliberately over time. A system may start by preparing work for approval. If evidence supports it, a specific class of action can later be permitted within defined limits. New conditions, recurring failures, or broken controls may justify reducing that authority again.

For leaders, this makes the assistant-or-agent question concrete. Choose one task and describe what the person currently does between receiving it and declaring it complete. Which parts are ready to be delegated, and which unresolved judgments would you otherwise be asking the system to make for you?

Enterprise AI Strategy: The Choices That Need to Fit Together

An enterprise AI strategy earns its value by making investment choices coherent.

Across my strategy and governance work, the recurring challenge has been connecting decisions that different groups make separately. A business team identifies an opportunity. Technology considers platforms. Security examines access. Finance asks about returns. People leaders plan learning. Each perspective is necessary, and their choices need to work together.

The starting point is the organization's strategy. What must improve for customers? Where does the business need to grow, become more resilient, or operate differently? Which capabilities make it distinctive? AI priorities should express those choices clearly enough that leaders can explain why one opportunity deserves resources ahead of another.

An objective such as improving customer retention creates a more useful conversation than increasing AI adoption. It opens several possible approaches: improving service, identifying recurring friction, helping employees make better decisions, or changing an offering. AI may contribute to any of them. The business outcome provides the basis for comparison.

It also creates room for both efficiency and opportunity. An enterprise needs to improve familiar work while examining how changing capabilities may alter customer expectations, competitive advantage, or the services it can offer. A portfolio that includes both can make these different ambitions visible and fund them appropriately.

The human-centered part belongs at the beginning. Understand the work as people experience it, including informal coordination, exceptions, and judgment that process diagrams omit. Ask who will benefit, whose workload may increase, and what employees need in order to use the new capability effectively.

Include the customer in that examination. A feature that is technically impressive may introduce friction into an experience people already value. The proposed improvement should be recognizable to its intended beneficiary.

Architecture provides the next set of choices. In my strategy outline, I separated systems of record, organizational context, and AI touchpoints. That remains a useful way to ask where information lives, which sources are authoritative, and how AI will interact with them.

An enterprise should be able to explain how a system gets the information it needs, how access follows the user's or agent's role, where outputs go, and how activity can be inspected. It also needs a way to keep business rules and context current. A capable model working with stale guidance can be consistently wrong in a very polished way.

Build, buy, and partner decisions follow from the capability required. Buying can be appropriate when an established product meets the need and fits the operating environment. Building can make sense where distinctive workflows, knowledge, or customer experience justify sustained ownership. A partner can provide expertise or capacity the organization needs to develop.

These choices include obligations after delivery. Who maintains the solution, handles incidents, evaluates changes, and owns the knowledge needed to operate it? A partner arrangement should be explicit about those responsibilities and about how the organization retains the ability to change direction.

Model and platform selection deserve the same discipline. Evaluate representative work, including difficult cases. Consider output quality, access controls, integration, response time, operating cost, observability, and the ability to recover from failure. The most capable model on a general benchmark may not provide the best fit for a particular business process.

A platform decision has wider consequences than a model decision. It can shape identity, information access, monitoring, development practices, and future switching costs. Leaders should know which elements they are standardizing and where variation is justified. Too much fragmentation creates duplication; premature uniformity can exclude useful approaches.

The pilot should resolve a decision. State what is uncertain, what evidence the test will produce, and what would justify expansion. That could involve performance on representative work, employee ability to use the system, integration reliability, or the economics of operating it at the intended volume.

A successful demonstration is only part of that evidence. Scaling requires an accountable business owner, service ownership, support capacity, appropriate data access, quality controls, and a credible benefit. It may also require changing roles and handoffs that the pilot bypassed through extraordinary attention from a small team.

Funding should reflect these differences. An exploratory experiment, a broadly available productivity tool, and a critical production workflow do not need identical business cases. They do need explicit purposes, cost visibility, and decisions about when to continue or stop.

The strategy should also preserve options. Document why a platform was selected, which assumptions underpin the business case, and which changes would trigger a review. Keep organizational knowledge usable beyond a single application where practical. The ability to revise a decision is valuable when both business needs and technology are evolving.

I would bring these choices together in one leadership conversation: the outcomes we seek, the people and workflows affected, the shared foundations we need, the responsibilities we will retain, and the evidence required to scale. Where does your current AI strategy make those choices explicit—and where are teams still making them independently?