
Are You Wasting Your Time With Harness Engineering?
Are explicit prompts, guardrails, and agent loops just patches for gaps in what the base model already knows?
117 people filled the AWS Experience at Floor28 for a live debate. Shuki Cohen (AI21) made the case for fine-tuning, Gad Benram (TensorOps) argued the harness side, and the audience kept both honest.
The harness case: most of the gains teams see in production now come from what surrounds the model, not the model itself. Context management, retrieval, tool design, evals, guardrails. Get those right and the model becomes a swappable part; every hour spent there keeps paying when the next generation ships.
The fine-tuning case: general models plateau on specialized work, and no amount of prompt plumbing buys back what owning the weights gives you. On narrow tasks with real data, a tuned model wins on consistency, latency, and unit cost, and it keeps winning at scale.
The two sides met in the middle more than either wanted to admit: start with the harness, earn the right to fine-tune with data and evals, and let cost decide the rest. Yaakov Tayeb (AWS) followed with how AWS builds these systems in practice, and the night closed with a panel Q&A and a live audience vote.





The next event is already on the calendar.
One short form makes you a member. The next event invite lands in your inbox, and the chats are waiting.
