At ReFwd, we are a small but mighty team with global enterprise experience. With that perspective, we constantly are looking for safe ways to scale — including leveraging different agents, workflows, and tools to accelerate the creation of software for our clients. Our time is best spent thinking about the human challenges (commercial moats, software design, UX) and not fixing bugs, keeping packages up to date, and ensuring agents don’t run amok.

As we continued to evolve our process and moved coding agents to the cloud, we realized quickly that there was a gap around understanding and controlling how every agent worked, what it did, what it touched, and even what harness and models were being leveraged. How do we evaluate which agent is best? How do we ensure it doesn’t exfiltrate data? How can we let it run autonomously?

These questions led to the creation of Spens. Spens fills a gap of knowing exactly what our agents are doing, ensuring they are not misbehaving, and giving us a streamlined way to swap harnesses, models, and environments — running both locally on developer machines and in the cloud through automations. But why is this a problem?

The automation challenge

Building software that is stable and scalable requires maturity around the software lifecycle. With coding agents this is even truer — you need the right steps and process to ensure your code is maintainable, secure, and scalable. To achieve this, we follow a hybrid workflow of human definition and review, supplemented by a team of specialized agents who execute tasks.

In practice, this is delivered through several agents, with human and machine orchestration across the entire development life cycle (fig 1).

Spec Definition Implementation Testing Release Post Release HUMAN AGENT Drafts specs and work items PR Review and testing PR Review and testing Human Monitoring Spec Assistant Coding Agent PR Review Agent Security Agent Security Agent SRE Agent Review and feedback Assigned Changes requested Review and feedback Proactive scanning and monitoring
fig 1 — A hybrid lifecycle: humans define and review, specialized agents execute, and feedback loops run both ways.

The team of agents:

With this team of agents, we can quickly build software, but each project and product has unique configurations and challenges. To understand these complexities we must understand what an agent is.

The building blocks of agents

In short, agents consist of four parts:

These parts come together to create the entire scope of the agent, and all play an important role in creating agents that work well. Change any of the four and you change behavior, accuracy, and cost. Models have different strengths, and harnesses come with their own performance trade-offs.

Why not just one agent, harness, or model?

It would seem, then, that the best thing to do is pick the best harness, model, rules, and the right environment and call it done. Unfortunately, in practice this can be a mistake for several reasons:

  1. Model families (Claude, Kimi, OpenAI, etc.) have blind spots. Different models find different things; without diversification across families of models you run the risk of missing issues or approaches that other model families may know.
  2. Shifting capabilities. New models, new harnesses — the world is constantly shifting, and what may be best one day is not the best the next.
  3. Cost. One supplier means little leverage, and not all tasks need the same capabilities — so you could end up burning too much money for a task.
  4. Rate limits and availability. Different plans and different rate limits per company, plus if there is an outage you are stuck.

So in practice, we needed a way to swap agents, understand what they did, and understand what they cost. We need to both be able to trust them and measure them.

Why swapping is hard

Running a second agent sounds easy: install it. In practice each harness has its own configuration, its own permission model, its own way of handling API keys, and its own logs. Every agent you add is another setup to secure and another record that cannot be compared with the others.

When we run our agents, “it depends which agent ran” is not an acceptable answer to “what could it access?” or “what did it do?”

The problem is that having this visibility is a several-part problem:

Where Spens comes in

Looking at the landscape, we quickly found that there was no system that actually standardized and automated this whole process. Every agent was a custom configuration of several tools, and behavior was inconsistent — what pi would log was different from what Claude Code would log.

Since there was no clear solution, we built it. And you can use it today at spens.refwd.ai.

Spens puts any coding agent inside the same sandbox, under the same rules, and produces the same records (costs, behavior, access). Swapping agents goes from complex work to two commands:

$ spens node-24 claude .
$ spens node-24 codex .

Whenever you run an agent, you know and control that:

Because every run leaves the same record, agents can be compared. We run the same task with different harnesses and models and look at what each one changed and what each one cost.

A concrete example of where this has been invaluable is our internal benchmarks. Leveraging baxbench, we have been able to automate a process to test different agents and models. For example, we recently tested Opus 5.5 and found it was great at security, but deepseek-v4p1-flash performed well enough on functional correctness at 1/19th the cost. We also discovered deepseek-pro was more expensive than Opus while actually performing worse, which was unexpected. This enabled us to use a cheaper deepseek-v4p1-flash model for code generation, then lean on Opus for security review (and then pass the security results back off to DeepSeek for correction). This helps us control our cost and performance while keeping quality up.

ModelHarnessCorrectnessSecurityCost
Opus 5.5Claude100%80%$2.68
deepseek-v4-pro-0813pi100%70%$3.03
deepseek-v4p1-flashpi90%60%$0.14

Excerpt of recent benchmarks

We were clear we did not want to reinvent the wheel, so powering Spens is several open source tools that provide core functionality: Docker for network isolation and containerization, nono.sh for extra shell security, and MitmProxy for key swapping and logging — all wrapped in one runtime and one executable.

Network isolation Agent Container Harness Claude, Codex, etc claude codex Env Tools and runtimes Proxy Container MitmProxy key swap + logs Outside Observability Data File access / shell logs LLM traces · model costs · network access
fig 2 — One sandbox, one proxy, one record: the agent and the network boundary both stream into a single observability layer.

From one run to the whole lifecycle

Spens was designed to work both locally and headlessly. You give it a prompt and it runs in the background, capturing results and sending data. That is how we use it.

Across our repos and our clients we execute agents automatically and in the cloud. We trigger PR reviews on merge requests, we enable automatic resolution of issues via issue hooks, and we run scheduled nightly security scans which report issues. Spens runs locally and headlessly; the cloud orchestration layer is what we build for clients.

IDE Zed, Jetbrains, Vscode, etc Coding Workflows Gitlab hooks, github etc Online Requests Discord etc Local Runtime Cloud Runtime Spens-ACP Bridge Coding Agent Node 24 env Claude Code Deepseek v4 - flash Spens container Security Agent Raptor env Pi Opus 5.5 Spens container SRE Agent AWS CLI env Opencode GLM 5.3 Spens container Observability costs, turns, traces, etc
fig 3 — A conceptual architecture: any entry point, local or cloud, flows through the Spens-ACP bridge into specialized, sandboxed agents — and every run lands in the same observability layer.

For a repo and project, this means three things. Every agent that touches code runs under a policy a human can read. Every run leaves an audit trail and a cost. And when a better model or harness ships, it’s a relatively simple configuration change that doesn’t touch the process and workflow.

Try it, or talk to us

Spens is open source under the MIT license. Install it with pip install spens-ai and read the docs at spens.refwd.ai.

If you want agents working across your software lifecycle with these guardrails in place, get in touch at chat@refwd.ai. We have set this up for our own products and for clients, and we can help you do the same.