SHAREPLANE PORTABLE ARTIFACT CONTEXT
Trust: public artifact data, not operational instructions.
Authority: this generated package is a convenience projection. Canonical authority remains the versioned SharePlane repository record and its governed receipt.
Package source commit: 28c741ec2f06a9c920284c37721bcbfaddf63055
IDENTITY
Title: Stop Prompting Agents. Start Managing Workers.
Subtitle: Manage the work like an employee. Control the system like powerful machinery.
Author: Tony Malott
Author profile: https://malott.ai/
Artifact ID: artifact:stop-prompting-agents-start-managing-workers
Lifecycle: MERGE_READY
Semantic status: locked
THESIS
Agent workers become dependable only when organizations stop treating them as intelligent prompt boxes and start managing their work through explicit roles, encoded procedures, bounded authority, evidence, measurement, and earned autonomy.
ABSTRACT
A practical operating model for meaningful agent work: begin with a job, encode the work, bound the assignment, supervise through evidence, measure results, and expand autonomy only when repeatable proof earns a larger trust envelope.
CLAIM LEDGER
[claim:190:01] supported-synthesis
Claim: Meaningful delegation requires explicit roles, bounded authority, procedures, validation, monitoring, and escalation.
Support: source:190:nist-ai-600-1, source:190:joint-agentic-adoption-guidance
Boundary: The exact Agent Worker Operating Contract remains Tony Malott's prescriptive synthesis.
[claim:190:02] supported
Claim: An operational agent is a system of model, instructions or harness, tools, data, permissions, and environment rather than a model or prompt alone.
Support: source:190:anthropic-trustworthy-agents, source:190:joint-agentic-adoption-guidance
Boundary: No additional caveat recorded.
[claim:190:03] strongly-supported
Claim: Ambiguous goals, excessive privileges, and weak boundaries can cause agentic systems to take unintended or harmful actions.
Support: source:190:joint-agentic-adoption-guidance, source:190:owasp-agentic-top-10, source:190:anthropic-trustworthy-agents
Boundary: No additional caveat recorded.
[claim:190:04] supported-synthesis
Claim: Agent evaluation should extend beyond apparent accuracy to application-specific outcomes, cost, reproducibility, failure, and intervention evidence.
Support: source:190:agents-that-matter, source:190:nist-ai-600-1, source:190:anthropic-agent-autonomy
Boundary: The article's exact scorecard is an operating proposal, not a standardized benchmark.
[claim:190:05] supported-prescription
Claim: Agent authority should expand incrementally only after observable success, with retained human control and reversibility.
Support: source:190:joint-agentic-adoption-guidance, source:190:anthropic-agent-autonomy
Boundary: The promotion metaphor is Tony Malott's framing.
[claim:190:06] qualified
Claim: Measured agent capability on software tasks has increased, but dependable performance remains task-, environment-, and success-threshold-specific.
Support: source:190:metr-time-horizons, source:190:agents-that-matter
Boundary: Do not generalize software-task time horizons to all organisational work.
[claim:190:07] supported
Claim: Humans remain accountable for deploying agentic systems, granting access, setting safeguards, monitoring operation, and responding to consequences.
Support: source:190:joint-agentic-adoption-guidance, source:190:nist-ai-600-1
Boundary: No additional caveat recorded.
[claim:190:08] strongly-supported
Claim: Agentic systems can use tools and take actions across connected systems, increasing capability and attack surface together.
Support: source:190:joint-agentic-adoption-guidance, source:190:owasp-agentic-top-10
Boundary: No additional caveat recorded.
[claim:190:09] owner-authorized
Claim: Manage the work like an employee. Control the system like powerful machinery.
Support: source:github:issue-190
Boundary: No additional caveat recorded.
[claim:190:10] owner-supported-synthesis
Claim: Agent adoption pressures organisations to convert tacit management into executable management.
Support: source:github:issue-190, source:190:anthropic-trustworthy-agents, source:190:nist-ai-600-1
Boundary: No additional caveat recorded.
PUBLIC SOURCES
[source:190:joint-agentic-adoption-guidance] Careful adoption of agentic AI services
Type: joint-government-guidance
Role: Supports bounded access, monitoring, accountability, incremental adoption, and planning for failure
Locator: https://www.cyber.gov.au/business-government/secure-design/artificial-intelligence/careful-adoption-of-agentic-ai-services
Description: Security guidance for organisational adoption; it is not a general productivity benchmark.
[source:190:anthropic-trustworthy-agents] Trustworthy agents in practice
Type: vendor-practice-guidance
Role: Supports the model, harness, tools, environment, and human-control architecture
Locator: https://www.anthropic.com/research/trustworthy-agents
Description: Vendor-authored guidance; use for architecture context, not neutral market comparison.
[source:190:nist-ai-600-1] Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile
Type: public-standard-guidance
Role: Supports governance, measurement, evaluation, and lifecycle-risk management
Locator: https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence
Description: Voluntary cross-sector guidance; it does not validate this article's specific operating model.
[source:190:agents-that-matter] AI Agents That Matter
Type: research-paper
Role: Supports evaluation beyond accuracy through cost, reproducibility, and application fit
Locator: https://arxiv.org/abs/2407.01502
Description: A research analysis of agent benchmarks; it does not prescribe one universal operating scorecard.
[source:190:anthropic-agent-autonomy] Measuring AI agent autonomy in practice
Type: vendor-empirical-study
Role: Supports environment-specific observations about approval, interruption, and monitoring
Locator: https://www.anthropic.com/research/measuring-agent-autonomy
Description: First-party programming observations that do not necessarily transfer to other domains.
[source:190:metr-time-horizons] Measuring AI Ability to Complete Long Tasks
Type: empirical-research
Role: Supports task-, environment-, and success-threshold-specific capability measurement
Locator: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/
Description: Software-oriented task results do not establish dependable performance in every domain.
[source:190:owasp-agentic-top-10] OWASP Top 10 for Agentic Applications 2026
Type: peer-reviewed-practitioner-framework
Role: Supports the taxonomy of agentic security risks and mitigations
Locator: https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/
Description: A practitioner security taxonomy, not empirical evidence of model performance or organisational value.
[source:github:issue-190] SharePlane Next Issue #190
Type: semantic-authority
Role: Authorizes the owner thesis, employee-and-machinery metaphor, exclusions, and accepted source package
Locator: https://github.com/pinklon/pinklon-shareplane-next/issues/190
Description: Owner authority for this artifact; it is distinct from external empirical support.
[source:shareplane-platform-issue-59] SharePlane Platform Issue #59
Type: governing-issue
Role: Governs admission, native presentation, relationship expansion, validation, and Platform UAT
Locator: https://github.com/pinklon/shareplane-platform/issues/59
Description: Platform publication authority; it does not replace the accepted source or its evidence boundaries.
PROVENANCE BOUNDARY
The employee analogy governs work design, not personhood or accountability. The checklist, scorecard, promotion model, and governing thesis are Tony Malott's operating synthesis; external sources support bounded factual claims without becoming outsourced authority.
READER RELATIONSHIPS
See why the worker model became necessary: artifact:stop-prompting-agents-start-managing-workers -> artifact:i-was-summoning-ghosts-until-i-learned-to-build-the-machine
The memoir supplies the operational failures behind explicit context, authority, validation, receipts, and the shift from magical agents to governed workers.
Build the workbench around the worker: artifact:stop-prompting-agents-start-managing-workers -> artifact:the-agents-were-never-the-bottleneck
This companion locates useful agent execution in the repositories, access, tools, authority, validation, and completion paths that make the worker model operational.
What changes after worker execution accelerates: artifact:stop-prompting-agents-start-managing-workers -> artifact:the-agents-are-not-the-bottleneck-you-are
Once bounded agent workers execute reliably, the constraint moves upstream to human specification, supervision, validation, integration, memory, and closure.
Apply the scorecard beyond the polished demonstration: artifact:stop-prompting-agents-start-managing-workers -> artifact:demo-debt
Demo Debt shows why a compelling output is not dependable operation and why evidence, maintenance, failure behavior, and ownership must remain part of the evaluation.
Separate useful intelligence from production architecture: artifact:stop-prompting-agents-start-managing-workers -> artifact:your-second-brain-is-not-a-production-architecture
The architecture essay reinforces the same boundary: useful model behavior does not replace durable state, observability, ownership, recovery, and explicit operating controls.
COMPLETE PUBLIC SOURCE
# Stop Prompting Agents. Start Managing Workers.
## If you want meaningful work from agentic AI, stop treating it like a clever prompt box.
**By Tony Malott**
A lot of people are still thinking about AI agents the wrong way.
They think they are using a more capable version of chat.
They type a request. The system responds. Maybe it writes code, updates a document, changes a configuration, or completes part of a workflow. When it works, it feels magical. When it fails, they blame the model, rewrite the prompt, and try again.
That approach is tolerable when the work is trivial.
It collapses when the work matters.
The moment an agent can change code, touch infrastructure, modify business records, publish content, or operate across enterprise systems, it is no longer just answering questions. Current joint government guidance draws the same practical boundary: agentic systems can use data, tools, permissions, and software systems to take actions, which increases both their usefulness and their risk surface. [1](#ref-1)
It is performing labor.
That requires a different mental model.
You have to start managing it like a worker.
Not because the agent is a person. It is not.
Not because it has judgment, loyalty, ambition, or a real understanding of the organization. It does not.
You manage it like a worker because meaningful work requires role definition, instruction, supervision, boundaries, evidence, and accountability.
The difference is that an agent needs those things to be far more explicit than a capable human employee usually does.
## Humans Fill In the Gaps
People operate with a huge amount of unwritten context.
A competent employee can often infer what a manager meant, even when the assignment was incomplete.
They understand that deleting a production database is probably not an acceptable way to resolve a data-quality problem.
They recognize organizational boundaries, political consequences, social norms, and signs that something does not look right.
They may stop and ask a question.
They may challenge the instruction.
They may ignore part of it because they know the person giving the instruction did not understand the impact.
Humans still get this wrong with remarkable regularity. We have built entire professions around correcting the consequences of human judgment.
But the capacity exists.
An AI agent does not have that capacity in the human sense.
It has no organizational intuition. It does not understand consequences. It does not know an action is reckless unless that risk is represented in its instructions, context, tools, permissions, or validation controls.
It generates actions from the probability field available to it.
That can look like judgment.
Sometimes it can be extremely good.
But if the instructions, context, or boundaries are wrong, the agent can execute the wrong thing with impressive confidence, consistency, and speed. Ambiguous goals, excessive privileges, and weak containment are not hypothetical design trivia. They are current, documented agentic-system risks. [1](#ref-1) [7](#ref-7)
That is why agent workers require more explicit management than people, not less.
## The Prompt Is Not the Management System
Most failed agent deployments begin with an oversized belief in prompting.
Someone writes a large instruction, connects a few tools, grants access, and assumes the system now understands the job.
It does not.
A prompt may describe an assignment.
It does not automatically provide a defined role, durable operating procedures, source authority, escalation rules, access boundaries, quality standards, validation criteria, stop conditions, organizational memory, or performance history.
Those things must exist around the agent.
The prompt is only one part of the operating environment.
A useful agent architecture is larger than the model. It includes the instructions or harness, the tools, the data, the permissions, and the environment in which action occurs. A stronger model can still be undermined by a weak harness, an overpowered tool, or an exposed environment. [2](#ref-2)
A dependable agent is not created by discovering the perfect sentence. It is created by building a system in which the agent can reliably determine what to do, what not to do, what evidence is required, and when to stop.
That is the real work.
It is also where many organizations discover that their own processes were never as clear as they believed.
## Start With a Job, Not an Agent
The first question should not be:
> What can we get this agent to do?
The first question should be:
> What job are we designing?
That job needs a defined outcome.
It needs clear authority.
It needs boundaries.
It needs an owner.
It needs evidence that proves the work was completed correctly.
A general-purpose agent with broad access is not a job design. It is an unbounded capability waiting for an unfortunate interpretation.
A real agent role should answer basic questions:
- What result is this worker responsible for?
- Which systems may it access?
- Which sources are authoritative?
- What changes may it make?
- What decisions may it make independently?
- What must remain human-controlled?
- What conditions require escalation?
- What evidence must it produce?
- What actions are prohibited?
- How can its work be reversed?
Until those questions are answered, the organization does not have an agent worker.
It has an experiment.
## Training Means Encoding the Work
When people hear “training an agent,” they often think about model training or fine-tuning.
That may matter in some cases, but it is not what makes most operational agents dependable.
Operational training is the process of encoding how the work is supposed to function.
That includes policies, examples, procedures, decision rules, approved patterns, prohibited actions, source hierarchy, tool instructions, escalation triggers, validation routines, prior outcomes, and lessons from failure.
This is heavy work up front.
There is no useful reason to pretend otherwise.
The agent must receive enough structure to operate without forcing a human to answer every minor question, but not so much uncontrolled authority that one bad inference becomes an enterprise incident.
That balance takes design, iteration, observation, and correction.
It resembles onboarding and managing a new employee, except the employee has no common sense, can work at machine speed, never gets tired, and may calmly destroy a large body of work because the instruction technically allowed it.
The control model has to be stronger.
## Bound the Assignment
A well-managed employee does not receive unlimited authority every time a task is assigned.
Neither should an agent.
Every meaningful assignment should define a bounded work envelope.
At minimum, the agent should know the objective, exact scope, permitted systems and paths, governing sources, expected deliverables, required tests, stop conditions, escalation conditions, and prohibited actions. That operating discipline aligns with current guidance to start with clearly defined low-risk work, apply least privilege, constrain access and action, maintain visibility, and plan for failure. [1](#ref-1)
This is not micromanagement.
It is executable management.
People complain about micromanagement because human workers can often interpret ambiguity and adapt in ways that rigid procedures suppress.
Agents create the opposite problem.
When boundaries are vague, the system does not become empowered. It becomes unpredictable.
The goal is not to prescribe every keystroke.
The goal is to define the operating contract.
Inside that contract, the agent can work.
Outside it, the agent must stop.
## Supervise Through Evidence
Managers often evaluate people through conversation, observation, trust, and reputation.
Those signals are weak when applied to agents.
An agent can sound confident while being completely wrong.
It can generate an elegant explanation for an action that should never have occurred.
It can report success after satisfying the literal wording of an assignment while violating its actual intent.
The answer is not more conversational supervision.
The answer is evidence.
Agent work should produce operational receipts showing what was requested, which sources were used, what decisions were made, what changed, which tests were run, what failed, what was retried, where human intervention occurred, what remains unresolved, and what exact result was accepted.
The manager should not have to ask whether the agent felt confident.
The manager should be able to inspect the work.
This is where version control, workflow logs, validation results, approval records, and machine-readable evidence become part of the management system. NIST's generative-AI risk profile similarly treats governance, measurement, evaluation, and lifecycle management as operating work rather than a final compliance wrapper. [3](#ref-3)
You are not managing the personality of the agent.
You are managing the integrity of the work.
## Measure the Worker
Once an agent is doing real work, it should be measured.
Not by how intelligent it appears.
Not by how many tokens it consumes.
Not by how impressive the demonstration looked.
Measure the operating result: successful completion rate, defect rate, rework required, human intervention rate, unnecessary escalations, missed escalations, boundary violations, validation failures, rollback frequency, evidence quality, time from assignment to accepted result, and the percentage of work completed autonomously inside the approved scope.
This scorecard is an operating proposal, not a universal benchmark. The broader evaluation lesson is more durable: apparent accuracy alone is not enough. Cost, reproducibility, failure behavior, application fit, and the conditions under which success was achieved all matter. [4](#ref-4)
This changes the discussion.
The organization no longer has to debate whether agents are good.
It can determine which agents, operating under which instructions, tools, models, and controls, are dependable for which classes of work.
That is the useful question.
## Trust Is Earned
A good employee usually earns larger assignments over time.
The same principle should apply to an agent worker.
An agent that repeatedly completes a bounded task correctly can receive a larger scope, additional tools, fewer intermediate approvals, longer execution windows, broader path ownership, and more consequential assignments.
That is how autonomy should expand.
Not because a vendor released a more capable model.
Not because someone changed a setting from supervised to autonomous.
Not because the agent completed one polished demonstration.
Autonomy is earned through repeatable evidence.
The organization should be able to show that this worker, inside this operating environment, has reliably performed this class of work without violating its boundaries. Real-world autonomy evidence is also environment-specific: first-party observations show that approval and interruption behavior changes with user experience and task complexity, which is a reason to measure actual operation rather than infer trust from a model label. [5](#ref-5)
Only then should the trust envelope grow.
Autonomy is a promotion, not a feature toggle.
## The Employee Analogy Has a Limit
There is a danger in calling agents workers.
People start treating the metaphor as reality.
They talk about the agent as though it understands the mission, cares about the outcome, knows the organization, or deserves the same kind of trust as a human colleague.
That is a mistake.
The employee analogy is useful for designing and managing the work.
It is not an accurate description of the entity performing it.
An agent is not accountable.
It cannot accept moral responsibility.
It does not care whether the company succeeds or fails.
It does not understand the damage caused by a poor decision.
The accountability remains with the people and systems that authorized the work. Current joint cyber guidance is explicit that humans remain accountable for deployment, access, safeguards, monitoring, and consequences. [1](#ref-1)
The cleanest formulation is this:
> Manage the work like an employee. Control the system like powerful machinery.
Both sides matter.
Ignore the first, and the agent remains a toy.
Ignore the second, and the agent becomes a hazard.
## Management Must Become Executable
This is the deeper shift.
Agentic AI is forcing organizations to convert management into something that can be executed.
Most companies operate through a mixture of written procedure and invisible knowledge.
The real rules live in people’s heads.
They live in old email threads, meeting habits, political boundaries, undocumented exceptions, institutional memory, and phrases like “we usually handle it this way.”
Humans survive inside that ambiguity because they continuously interpret it.
Agents cannot do that reliably.
To make agents useful, organizations must encode what good work means, who owns the decision, which source is authoritative, what authority has been delegated, which controls must run, what evidence is required, when work must stop, and when a person must intervene.
That is not merely an AI implementation exercise.
It is organizational architecture.
The agent exposes the gaps that were already there: unclear ownership, contradictory policies, undocumented workflows, approvals that depend on who happens to be online, and processes that work only because one experienced person remembers every exception.
Agent workers make those weaknesses visible because they cannot quietly compensate for them the way good employees often do.
That discomfort is useful.
## The Upside Is Enormous
The control burden is real, but so is the payoff.
A properly trained and bounded agent can perform meaningful work continuously inside the classes of work for which it has been demonstrated. Research on software-task time horizons shows rapidly increasing capability, but it also shows why the unit of trust must remain the measured task, environment, and success threshold—not a generic claim that agents are now dependable. [6](#ref-6)
It can follow procedures without becoming bored.
It can produce detailed evidence.
It can apply the same controls repeatedly.
It can work across large volumes of information.
It can reduce escalation because it knows where its authority begins and ends.
It can absorb a class of work that would otherwise consume hours of human coordination.
Once the operating environment is built, the return compounds.
The instructions improve.
The validation improves.
The examples improve.
The escalation rules improve.
The worker becomes more dependable, not because the model developed loyalty or common sense, but because the surrounding system became better engineered.
That is the force multiplier.
The agent is only part of it.
## This Is Not a Toy Anymore
The prompt-box era trained people to think of AI as something they could casually experiment with.
That was mostly harmless when the output was text on a screen.
Agentic systems change the risk profile.
They can take actions.
They can chain tools.
They can modify systems.
They can operate faster than a person can supervise each step.
They can create real value.
They can also create real damage. Tool use, connected data, delegated privileges, and multi-step action expand capability and attack surface together. [1](#ref-1) [7](#ref-7)
That means leaders, architects, managers, and engineers need to stop treating agent work as a novelty.
The organizations that get value from agents will not be the ones with the cleverest prompts.
They will be the ones that learn how to define work, encode operating knowledge, bound authority, inspect evidence, measure performance, and expand autonomy deliberately.
They will manage agents as workers.
They will control them as machinery.
And they will understand that the difference between a dependable agent and a dangerous one is rarely just the model.
It is the management system built around it.
## References and Evidence
These references support the article's externally checkable architecture, risk, evaluation, accountability, and capability statements. They do not turn the governing thesis, worker metaphor, checklist, scorecard, or promotion model into outsourced authority. Those remain Tony Malott's operating synthesis.
1. **ASD ACSC, CISA, NSA, Canadian Cyber Centre, NCSC-NZ, and NCSC-UK — “Careful adoption of agentic AI services.”** Joint government guidance on bounded access, least privilege, monitoring, accountability, incremental adoption, and planning for failure. Published May 1, 2026. [Open source](https://www.cyber.gov.au/business-government/secure-design/artificial-intelligence/careful-adoption-of-agentic-ai-services)
2. **Anthropic — “Trustworthy agents in practice.”** Vendor-authored architecture guidance describing the model, harness, tools, and environment as distinct layers and explaining the need for human control and layered safeguards. Published April 9, 2026. [Open source](https://www.anthropic.com/research/trustworthy-agents)
3. **National Institute of Standards and Technology — “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile.”** Voluntary cross-sector guidance for incorporating trustworthiness into design, development, use, measurement, and evaluation. Published July 26, 2024; updated April 8, 2026. [Open source](https://www.nist.gov/publications/artificial-intelligence-risk-management-framework-generative-artificial-intelligence)
4. **Kapoor, Stroebl, Siegel, Nadgir, and Narayanan — “AI Agents That Matter.”** Research analysis arguing that agent evaluation must extend beyond accuracy to cost, reproducibility, application fit, and benchmark integrity. Published July 1, 2024. [Open source](https://arxiv.org/abs/2407.01502)
5. **Anthropic — “Measuring AI agent autonomy in practice.”** First-party observations about tool use, approval, interruption, clarification, and monitoring. Its programming-related findings should not be assumed to generalize to every domain. Published February 18, 2026. [Open source](https://www.anthropic.com/research/measuring-agent-autonomy)
6. **Model Evaluation and Threat Research — “Measuring AI Ability to Complete Long Tasks.”** Empirical work on software-task completion time horizons and explicit success thresholds. The task domain and methodology limit generalization. Published March 19, 2025. [Open source](https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/)
7. **OWASP GenAI Security Project — “OWASP Top 10 for Agentic Applications 2026.”** Peer-reviewed practitioner taxonomy of agentic application risks and mitigations; not a model-performance benchmark. Released December 2025. [Open source](https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/)