
Your R&D Bottleneck Has Moved: How to Get More R&D Output from People and Tokens
An agent finishes a ticket overnight. The tests pass, the pull request is ready, and an engineer reviews it in the morning. What made that work: the model, the context it inherited, the way the task was delegated, or the infrastructure running underneath?
At our September 9 session, Your R&D Bottleneck Has Moved: How to Get More R&D Output from People and Tokens, we explored how those pieces fit together. Hosted by Cláudio Lemos and moderated by Gad Benram, CTO of TensorOps, the event brought together Renen Avneri, AI Architect at Riverside, and Elad Guttel, Co-Founder and VP of R&D at Valar.
The discussion connected two questions engineering leaders increasingly face: how to give agents useful autonomy, and how to pay for the work they do.
Give agents organizational memory
Renen opened with a question about value: how much useful work do we get for the attention we spend, whether that attention comes from a person or a model?
His examples came from Riverside’s ATA platform, short for “assistive to autonomous.” Agents begin alongside people, learning the product, the team’s conventions, and the decisions behind the code. As their operators build confidence in them, they can take on more independent work.
Gad compared the process to onboarding a junior engineer: teach them the system, help them build experience, then expand their responsibilities. That analogy made an architectural point concrete. An agent needs more than a task description; it needs a way to retain what the organization has already taught it.
Renen described agents with defined roles, reusable skills, and persistent memory. Each agent keeps its own working knowledge while drawing on shared team and organizational guidance. A debugging agent can remember where a particular control lives. An architecture agent can recall an earlier decision and who worked on it.
Those memories can then support work across developers’ machines and cloud environments. The value compounds when the next person benefits from what the agent learned with the previous one.
Make delegation visible
Riverside’s examples extended across the software development lifecycle: architecture reviews during planning, specialist code reviewers, testing agents, and support triage.
One example captured the practical benefit. An overnight support message could trigger an agent to investigate whether an issue was widespread and needed escalation, or whether it could prepare an answer for the support team. The useful outcome was fewer unnecessary interruptions for people.
That level of delegation depends on visibility. Renen showed how Agent Trail connects the person or event that triggered work to the agents, models, tools, and costs involved. Teams can see where work goes and what it costs.
Permissions remained specific to the job. He described a DevOps agent that needed an extended period of guidance before receiving its own runtime and credentials. Code review and QA agents had different requirements. Readiness depended on the responsibility being delegated.
Renen reported an 80% reduction in agentic platform costs through the practices he described. The Riverside platform video also reported a 40% reduction in time from commit to merge. These were results presented from Riverside’s deployment, with its particular workflows and infrastructure.
Context is an engineering asset
One of the most useful threads in Renen’s talk was the connection between agent context and organizational knowledge.
Repeatedly explaining the same system wastes human attention and model tokens. Summarizing a decision without preserving its reasoning makes future work harder. Handing someone an enormous document can delay them when a few relevant references would let them begin.
His practical starting point was to build the knowledge agents need: team guidance, reliable sources of truth, and records of architectural decisions. Keep that information current and make it accessible at the point of work.
Longer-running agents make this investment more valuable. Every handoff is an opportunity to preserve context or lose it.
Price inference around the task’s deadline
Elad approached the same problem from the infrastructure side. His opening question was whether teams were paying interactive prices for inference that nobody was interacting with.
Consider an agent fixing a flaky test overnight. If a person will review the result at 9 a.m., finishing a few minutes earlier may add little value. That flexibility creates room to batch requests, schedule work differently, and use compute more efficiently.
Elad described five sources of savings in Valar’s approach:
- Compute choice: using different accelerators where they suit the workload.
- Scheduling: taking advantage of batching and flexible timing.
- Cluster configuration: tuning serving capacity for the workload’s latency needs.
- Serving efficiency: improving cache reuse and reducing unnecessary context.
- Model selection: routing requests according to the capability they require.
He also explained why moving between models needs careful engineering. API differences, cache behavior, and missing capabilities can undermine the savings. Difficult requests still need access to stronger models.
Gad brought the discussion back to a concrete example: a Mario-style game he had generated using Valar behind his coding environment. He reported that the environment displayed an estimated cost of about $10.50, while the actual Valar charge was about $1. It was one illustrative run, showing why the cost of completed work deserves attention alongside token prices.
Measure the outcome, then expand
Latency still matters when a person is waiting or one agent’s output blocks the next step. Elad’s proposed “night shift” makes urgency part of the task: complete this work by a deadline, then allocate resources around that requirement.
Evaluation matters just as much. In the Q&A, he explained that Valar’s routing work was strongest on coding tasks, where tests and other checks help assess success. Work with less easily verifiable outcomes is harder to evaluate.
Together, the talks suggested a practical starting point: choose a repeatable workflow, give the agent the context and access it needs, define how success will be checked, and measure both cost and human attention. Expand responsibility as the evidence supports it.
The session closed with an announcement of a community offer: $100 in Valar credits and a conversation about architecture and getting started.
Thanks to Renen, Elad, Gad, Cláudio, and everyone who joined with questions. Explore upcoming Inference Hub events and bring your own experience to the next conversation.

The next event is already on the calendar.
One short form makes you a member. The next event invite lands in your inbox, and the chats are waiting.


