
Long Horizon AI and Context Extension in LLMs
Agents don't just forget facts. They forget why.
Two talks on what changes when an AI system is expected to keep working long after the prompt is over.
Gad Benram (TensorOps) opened with the shift from seconds-long completions to long-horizon AI: agents that run for hours, days, or months toward an objective, taking autonomous actions, researching, and using tools through a fog of war, without a human approving every step. The bottleneck is memory. Models are stateless, so every turn is a conversation with someone who has no short-term memory, and once the context window fills mid-task the agent forgets catastrophically.
Ori Siegal picked up from there and framed it as an alignment problem, with the Cobra Effect in colonial India as the cautionary tale: an incentive that looked sensible produced the opposite of what anyone wanted. Context works the same way. He showed how adding distracting information to a prompt can drop a model's accuracy from 98% to 34%, which rules out dumping raw data into the context as a strategy.
His answer was a four-step method for building a second brain for an agent system:
- Capture. Find the operational bottlenecks and pull raw knowledge out of databases, APIs, and conversations with people.
- Store. Put it into a structure the agent can use: vector stores, temporal graphs, or plain, clean Markdown.
- Maintain. Keep it alive. Refresh stale facts, resolve contradictions, and track the things that shift underneath you, like an API that changed.
- Enforce. Make sure the agent actually uses it, through strict prompting, skill checks, and dedicated librarian agents.
Both speakers landed in the same place: the work is moving away from training foundation models and toward harness engineering. The next few years belong to the teams building the environments and ontologies where fleets of agents coordinate, manage their own context limits, and get work done alongside people.
- KEYNOTE
From Minutes to Days: The Rise of Long-Horizon AI
AI is moving from answering questions to doing work: agents that operate for hours or days, coordinate across steps, use tools, recover from failures, and keep working toward a goal with limited human involvement. The techniques behind the shift — harness engineering, tool use, state and memory, verification, feedback loops, multi-agent collaboration — and the use cases that open up as the time horizon expands.
- TALK
Agents don't just forget facts. They forget WHY.
Long-horizon agents rarely fail because the model gets worse at step 400; they fail because step 400 can no longer see why step 12 made a decision. Memory as a governance problem rather than a storage problem: what knowledge is worth preserving and in what form, what to discard early, how to tell whether information is still valid, and what makes a rule actually influence an agent when it matters — from four months of building and operating such a system daily.
The next event is already on the calendar.
One short form makes you a member. The next event invite lands in your inbox, and the chats are waiting.

-300x300.jpg&w=128&q=75)


