Software used to wait to be told what to do. It took input, applied logic, and returned output. Agents changed that. They decide, they act, they touch systems, and increasingly they act without a human standing over them. That shift — from software that computes to software that chooses — is the moment the ethics of engineering stopped being a compliance exercise and became an operational one.
Agency is the new interface
A function call is deterministic. Given the same input it does the same thing, and if it is wrong, the bug is in the code. An agent is different. Given the same input it may reason its way to a different plan, call a tool you did not anticipate, or take an action you did not explicitly authorize. The blast radius is no longer a stack trace. It is a deleted table, a leaked credential, a customer who was told the wrong thing.
We do not get to treat that as a purely technical problem. When a system can make decisions that affect people’s money, health, employment, or privacy, the people who build it carry a moral obligation to make those decisions legible, bounded, and accountable. Capability without accountability is not innovation. It is negligence with better marketing.
The accountability gap
The uncomfortable truth about most agent deployments is that nobody can answer the simplest question: what did it actually do? Teams ship an agent, wire up a few tools, and cross their fingers. When something goes wrong, the investigation is archaeology — scattered logs, missing prompts, an API key nobody remembers rotating. “It depends which agent ran” is not an answer you can give a regulator, a customer, or your own board.
That gap is where ethics and engineering meet. You cannot be responsible for a system you cannot observe. You cannot consent to a risk you cannot measure. And you cannot fix a failure you cannot reproduce.
An agent that cannot be observed cannot be trusted. An agent that cannot be trusted has no business acting on its own.
Three imperatives
Every team putting agents into production should be able to say three things without hesitation.
- Observability. Every LLM call, tool call, network request, and file change is recorded and attributable. Not for voyeurism — for accountability. If you cannot reconstruct what an agent did and why, you are not operating it. You are hoping.
- Containment. Agents run with least privilege. They touch the files you intend, reach the network you allow, and never see the real secrets — placeholders go in, and a proxy swaps them at the boundary. The default is no, and access is granted deliberately.
- Accountability. A named human owns every autonomous run. There is a policy a person can read, a cost a person can justify, and a rollback a person can trigger. Autonomy is a privilege you grant, not a switch you forget.
The ethics of cost and access
There is a quieter ethical dimension to model choice. Routing every task to the largest frontier model is not just expensive — it concentrates capability, cost, and risk in a handful of vendors. Diversifying across model families is a resilience strategy, but it is also a fairness one. It keeps a single provider’s blind spots from becoming your users’ problem, and it keeps the economics of intelligence from becoming a moat that only the largest enterprises can cross.
The same logic applies to data. An agent that exfiltrates context to a third party is a data-processing decision whether or not anyone wrote it down. The ethical default is the conservative one: data stays inside your walls, access is explicit, and every transfer is logged.
Where humans belong
None of this argues for keeping humans in every loop. A human rubber-stamping every action is theater, not oversight, and it destroys the very leverage that makes agents worth building. The point is to put people where judgment actually matters — defining intent, setting policy, reviewing outcomes, and owning consequences — and to let machines handle the repetitive work in between.
That is the division of labor we should want: humans decide what should happen, agents decide how, and the system makes sure we can always tell the difference.
The imperative
Agents are going to run more of the software that runs our lives. The only question is whether they do it inside a system that can see them, bound them, and answer for them — or in the dark. That is not a feature you add at the end. It is the foundation you build on from the first commit.
At ReFwd we build agents the way we would want them built for us: observable by default, contained by policy, and owned by a human who can explain every decision. Not because it is easy, and not because a regulation demands it, but because it is the only version of this future worth shipping.
