At ReFwd, we are a small but mighty team with global enterprise experience. With that perspective, we constantly are looking for safe ways to scale — including leveraging different agents, workflows, and tools to accelerate the creation of software for our clients. Our time is best spent thinking about the human challenges (commercial moats, software design, UX) and not fixing bugs, keeping packages up to date, and ensuring agents don’t run amok.
As we continued to evolve our process and moved coding agents to the cloud, we realized quickly that there was a gap around understanding and controlling how every agent worked, what it did, what it touched, and even what harness and models were being leveraged. How do we evaluate which agent is best? How do we ensure it doesn’t exfiltrate data? How can we let it run autonomously?
These questions led to the creation of Spens. Spens fills a gap of knowing exactly what our agents are doing, ensuring they are not misbehaving, and giving us a streamlined way to swap harnesses, models, and environments — running both locally on developer machines and in the cloud through automations. But why is this a problem?
The automation challenge
Building software that is stable and scalable requires maturity around the software lifecycle. With coding agents this is even truer — you need the right steps and process to ensure your code is maintainable, secure, and scalable. To achieve this, we follow a hybrid workflow of human definition and review, supplemented by a team of specialized agents who execute tasks.
In practice, this is delivered through several agents, with human and machine orchestration across the entire development life cycle (fig 1).
The team of agents:
- Spec Assistant — helps with creation and writing of specifications.
- Coding Agent — agents which specifically write code.
- PR Review Agents — agents specialized in reviewing PRs.
- Security Agent — agents focused on interrogating security positions.
- SRE Agents — agents focused on SRE tasks.
With this team of agents, we can quickly build software, but each project and product has unique configurations and challenges. To understand these complexities we must understand what an agent is.
The building blocks of agents
In short, agents consist of four parts:
- The harness is the runtime, such as Claude Code, Codex, or pi. It runs the loop with the model, exposes the tools, and coordinates the work.
- The model is the LLM underneath.
- Skills and rules are the configuration that tells the agent how to do a particular job.
- The environment is the machine, tools, and access the agent runs with.
These parts come together to create the entire scope of the agent, and all play an important role in creating agents that work well. Change any of the four and you change behavior, accuracy, and cost. Models have different strengths, and harnesses come with their own performance trade-offs.
Why not just one agent, harness, or model?
It would seem, then, that the best thing to do is pick the best harness, model, rules, and the right environment and call it done. Unfortunately, in practice this can be a mistake for several reasons:
- Model families (Claude, Kimi, OpenAI, etc.) have blind spots. Different models find different things; without diversification across families of models you run the risk of missing issues or approaches that other model families may know.
- Shifting capabilities. New models, new harnesses — the world is constantly shifting, and what may be best one day is not the best the next.
- Cost. One supplier means little leverage, and not all tasks need the same capabilities — so you could end up burning too much money for a task.
- Rate limits and availability. Different plans and different rate limits per company, plus if there is an outage you are stuck.
So in practice, we needed a way to swap agents, understand what they did, and understand what they cost. We need to both be able to trust them and measure them.
Why swapping is hard
Running a second agent sounds easy: install it. In practice each harness has its own configuration, its own permission model, its own way of handling API keys, and its own logs. Every agent you add is another setup to secure and another record that cannot be compared with the others.
When we run our agents, “it depends which agent ran” is not an acceptable answer to “what could it access?” or “what did it do?”
The problem is that having this visibility is a several-part problem:
- File / access sandboxing — ensuring the agent doesn’t touch files it shouldn’t (nono.sh, microVMs, Daytona).
- Network sandboxing — ensuring it doesn’t make requests it shouldn’t, and keeping credentials out of the agent (agent proxy).
- Observability — knowing what tool calls and LLM calls were made (Langfuse, Arize Phoenix).
Where Spens comes in
Looking at the landscape, we quickly found that there was no system that actually standardized and automated this whole process. Every agent was a custom configuration of several tools, and behavior was inconsistent — what pi would log was different from what Claude Code would log.
Since there was no clear solution, we built it. And you can use it today at spens.refwd.ai.
Spens puts any coding agent inside the same sandbox, under the same rules, and produces the same records (costs, behavior, access). Swapping agents goes from complex work to two commands:
$ spens node-24 claude .
$ spens node-24 codex .
Whenever you run an agent, you know and control that:
- It’s sandboxed — the agent only touches the files and tools you want.
- It’s contained — you set the network access rules and what resources it can reach.
- It doesn’t know your secrets — the agent sees placeholders, and the network proxy swaps the real keys in.
- You will have a log — everything is recorded: every LLM call, tool call, HTTP request, and file change.
- You decide what stays and what goes, with rollback and restoring capabilities.
Because every run leaves the same record, agents can be compared. We run the same task with different harnesses and models and look at what each one changed and what each one cost.
A concrete example of where this has been invaluable is our internal benchmarks. Leveraging baxbench, we have been able to automate a process to test different agents and models. For example, we recently tested Opus 5.5 and found it was great at security, but deepseek-v4p1-flash performed well enough on functional correctness at 1/19th the cost. We also discovered deepseek-pro was more expensive than Opus while actually performing worse, which was unexpected. This enabled us to use a cheaper deepseek-v4p1-flash model for code generation, then lean on Opus for security review (and then pass the security results back off to DeepSeek for correction). This helps us control our cost and performance while keeping quality up.
| Model | Harness | Correctness | Security | Cost |
|---|---|---|---|---|
| Opus 5.5 | Claude | 100% | 80% | $2.68 |
| deepseek-v4-pro-0813 | pi | 100% | 70% | $3.03 |
| deepseek-v4p1-flash | pi | 90% | 60% | $0.14 |
Excerpt of recent benchmarks
We were clear we did not want to reinvent the wheel, so powering Spens is several open source tools that provide core functionality: Docker for network isolation and containerization, nono.sh for extra shell security, and MitmProxy for key swapping and logging — all wrapped in one runtime and one executable.
From one run to the whole lifecycle
Spens was designed to work both locally and headlessly. You give it a prompt and it runs in the background, capturing results and sending data. That is how we use it.
Across our repos and our clients we execute agents automatically and in the cloud. We trigger PR reviews on merge requests, we enable automatic resolution of issues via issue hooks, and we run scheduled nightly security scans which report issues. Spens runs locally and headlessly; the cloud orchestration layer is what we build for clients.
For a repo and project, this means three things. Every agent that touches code runs under a policy a human can read. Every run leaves an audit trail and a cost. And when a better model or harness ships, it’s a relatively simple configuration change that doesn’t touch the process and workflow.
Try it, or talk to us
Spens is open source under the MIT license. Install it with pip install spens-ai and read the docs at spens.refwd.ai.
If you want agents working across your software lifecycle with these guardrails in place, get in touch at chat@refwd.ai. We have set this up for our own products and for clients, and we can help you do the same.
