Today we're releasing Context-Bench, an open benchmark for agentic context engineering. Context-Bench evaluates how well language models can chain file operations, trace entity relationships, and manage long-horizon multi-step tool calling.
29 likes 1 replies
?