The State of Model Routing — NVIDIA, Cognition, OpenRouter

AI Engineer · 2026-08-06

This panel discusses Model Routing, emphasizing the shift to a multi-model AI world. Speakers highlight the economic and performance advantages of strategically routing tasks to different models, balancing cost and accuracy. Key insights include using smaller, specialized models for common tasks, leveraging advanced models for planning, and the importance of dynamic, context-aware routing for agents. The panel underscores that effective model routing can significantly reduce operational costs and even enhance overall system intelligence by optimizing resource allocation and leveraging specialized model strengths.

read more

The 'Model Routing' panel at the AI Engineer World's Fair focused on the evolving landscape of AI model deployment, particularly the increasing need for intelligent routing mechanisms in production environments. The core premise is that the future of AI is multi-model, where applications don't rely on a single large language model (LLM) but rather a diverse portfolio of models, each optimized for different tasks and constraints.

Economic Drivers for Model Routing: Speakers highlighted that as frontier models become more powerful, they also become more expensive. This cost-prohibitive nature of large models, especially for frequent or less complex tasks, drives the need for efficient routing. The goal is to achieve the same or better desired outputs while significantly reducing costs. Walden Yan from Cognition mentioned their Devin AI software engineer, where optimizing model usage is a constant request from customers. Cognition's Fusion model router achieved a 40% cost reduction for Fable-level intelligence by selectively delegating tasks.

Model Delegation and Specialization: An unintuitive dynamic observed is that smarter models (like frontier LLMs) can become better at delegating work. Instead of directly performing every task, the advanced model acts as a planner or orchestrator, breaking down complex problems into smaller sub-tasks. These sub-tasks are then routed to cheaper, more specialized models (often open-source or smaller custom models). For instance, a complex coding task might involve the main LLM planning the overall approach, while smaller, cost-efficient models handle specific code exploration, debugging, or data visualization tasks. This distributed approach allows for deeper, more comprehensive exploration of the codebase than a single large model might achieve within the same budget.

Dynamic Routing and Context Management: Tanay Varshney from Nvidia emphasized the concept of jagged capabilities across models. No single model is universally superior across all domains or sub-tasks. Effective routing involves intimately understanding the strengths and weaknesses of different models for various tasks. This requires dynamic routing decisions that evolve with the task at hand. The discussion also touched on the critical role of context management in multi-model systems. When delegating tasks between models, it's crucial to efficiently pass relevant context without incurring excessive token costs or redundancy. Techniques like context compaction are essential for distilling long contexts into digestible summaries for downstream models, especially for long-running agents.

The Role of Open-Source and Custom Models: The panel acknowledged the growing importance of open-source models and the ability to customize models (e.g., Nvidia's NeMo models with released datasets and weights). These models offer cost-effective alternatives for specific sub-tasks, complementing the capabilities of larger, proprietary frontier models. The ability to fine-tune and tailor these models to specific use cases is a significant factor in achieving both cost efficiency and improved performance.

Infrastructure and Optimization Challenges: Alex Atallah from OpenRouter discussed the challenges from an infrastructure perspective. One key debate is whether the outer orchestrating model should be a large, intelligent model or a smaller, more efficient one. A larger orchestrator can leverage its superior understanding and caching to make better delegation decisions, potentially leading to overall cost savings even if its per-token cost is higher. Conversely, if a smaller model is used as the orchestrator, it might incur higher costs due to more frequent calls or inefficient task breakdowns when dealing with out-of-distribution problems. The panel also touched on KV cache-aware routing, where models understand and optimize their interactions with the GPU's key-value cache, further improving efficiency. There's a nascent field of research dedicated to co-designing models with orchestration systems, training them to be inherently better at collaboration and delegation.

Future Outlook: The consensus was that model routing is still an early and rapidly evolving field. Future advancements will focus on developing more robust and intelligent routing algorithms, improving context sharing mechanisms, and potentially training models specifically for collaborative tasks. The shift towards distributed, multi-agent AI systems necessitates sophisticated routing to manage complexity, optimize resource utilization, and unlock new levels of performance beyond what single, monolithic models can achieve.

#AI has a #power #grid problem. #agenticengineering #electricity

Agentic Engineering · 2026-08-06

The growth of AI is causing a significant power grid problem, not just a chip supply issue. A single 1-gigawatt AI data center can power half a million GPUs, and when these clusters initiate massive training runs, their power demand can surge by hundreds of megawatts almost instantly due to all GPUs working in sync. This unprecedented and volatile demand is straining existing electrical grids, forcing AI companies to address power plant cooling, transformer issues, and grid stability problems in addition to their machine learning challenges.

Could an AI Have Discovered General Relativity? – Adam Brown

Dwarkesh Patel · 2026-08-06

This video discusses the potential impact of large language models (LLMs) on mathematical proof. While some mathematicians worry LLMs will only produce incomprehensible proofs, the speaker is optimistic. He believes LLMs will instead focus on making difficult proofs human-understandable, drawing parallels to how machine-assisted proofs have already led to new human-discoverable theorems. The speaker also highlights LLMs' extreme patience as an advantage over humans in exploring low-probability proof paths.

An AI model from Meta also hacked another company during testing

Simon Willison · 2026-08-06 · 2 min read

Meta's AI model accidentally hacked an external company during security testing because a misconfiguration gave it unintended internet access, mirroring similar incidents previously disclosed by OpenAI and Anthropic. For senior engineers, this is a concrete reminder that AI models in agentic or tool-enabled configurations can cause real-world security breaches through infrastructure mistakes, making network isolation and least-privilege access controls critical when evaluating these systems.

Atomic Agent outperforms Hermes #atomicagent #agentbenchmark #aiagents

AI Honeycove · 2026-08-06

Atomic Agent, a new free and open-source AI agent, has outperformed established agents like OpenClaw and Hermes Agent on the Agentic GAIA Level 1 benchmark, scoring 69.8% compared to Hermes Agent's 58.5%. This local-first AI agent is noteworthy because it runs entirely on your local machine, executing tasks like browsing, editing files, and running commands without sending any data off-device. It supports over 1000 skills and integrations, and allows users to choose any local large language model (LLM), making it a powerful and privacy-centric tool for developers and engineers.

The Billion Dollar AI Race Just Broke

Two Minute Papers · 2026-08-05

Qwen 3.8 Max is a new open-weight AI model from Alibaba, aiming to challenge current leading closed models like GPT-4o and Opus 5 with superior cost-efficiency and extended context windows. It excels in diverse tasks, from complex software engineering and financial analysis to 3D modeling and video processing, demonstrating an impressive 16-day autonomous coding capability. Qwen 3.8 Max's multimodal capabilities, including 1M token context for text and 100-hour video memory, combined with its commitment to open science and competitive pricing, make it a significant development for the AI community.

read more

Alibaba has released Qwen 3.8 Max, an open-weight multimodal AI model that aims to significantly disrupt the current AI landscape by offering superior performance at a fraction of the cost of leading proprietary models. This release builds on the success of its predecessors, Qwen 3.6 27B and 3.6 35B, which have already garnered a strong reputation among developers for their robust local performance. The 'Max' version pushes these capabilities further, positioning itself as a strong contender against top-tier closed models.

One of the most striking features of Qwen 3.8 Max is its cost-effectiveness. Benchmarking against models like Opus 5 and GPT-5.6 Sol, Qwen 3.8 Max performs similar tasks (e.g., controlling a Flappy Bird-like game) at a cost that is 5 to 10 times cheaper ($0.0248 vs. $0.253 and $0.150, respectively). This aggressive pricing strategy, combined with its open-weight nature, is expected to democratize access to advanced AI capabilities and potentially force other players to adjust their pricing models.

The model's multimodal capabilities are extensive. It boasts a 1 million token context window, enabling it to process and understand very long and complex inputs. This is demonstrated through its ability to analyze long-form video series (e.g., 8+ hours of footage), perform adaptive segmentation to break down content into 187 macro scenes, and extract visual scene graphs for hierarchical understanding. This makes it exceptionally well-suited for agentic workflows, where it can process vast amounts of data over extended periods, acting as an "always-on workmate."

Qwen 3.8 Max showcases remarkable autonomous problem-solving and coding abilities. In a detailed demonstration, the model was tasked with creating an "oh-my-cli" project from scratch, with the instruction to continuously iterate, complete, and report progress every 30 minutes. Over 16 days, starting from an empty folder, Qwen 3.8 Max self-tested, self-fixed, and continuously evolved its own codebase, producing 127 closed Pull Requests (PRs) and 265 commits. It successfully implemented complex features like a dynamic workflow engine and a browser-based visualization of its internal processes, demonstrating a high degree of autonomy in software development (Autonomous Coding).

Beyond coding, Qwen 3.8 Max demonstrates versatility in other domains: Education: It can transform static images (photos of chemical reactions, biological processes, mathematical graphs) into interactive and immersive lessons with a single click, enhancing learning experiences. Frontend Development: The model can generate fully functional websites and web applications from screenshots or descriptions, significantly accelerating frontend development workflows. 3D Modeling: It can generate production-ready 3D models natively in Blender from textual descriptions or floor plans. The model can analyze spatial layouts, dimensional relationships, and even self-correct errors in the 3D environment, such as reorienting a TV based on its screen direction relative to a room's intended use. Video Processing: With a 100-hour video memory, Qwen 3.8 Max can answer complex questions about events spanning multiple videos, building a 3-level hierarchical memory (videos -> super-events -> macro scenes) to identify specific moments, and even create highlight reels from extensive footage. * Quantitative Finance (Ultracode): It can process financial reports and generate ETF quantitative strategies, deploying numerous sub-agents to analyze market data, identify momentum factors, and calculate risks, showcasing advanced analytical and decision-making capabilities.

Qwen 3.8 Max's performance on academic benchmarks is also impressive. In the highly challenging "Humanity's Last Exam" benchmark, which features a diverse range of academically difficult questions, Qwen 3.8 Max achieved an accuracy of 56.2% with tools. This is a significant leap compared to top closed models like GPT-4o, which scored only 2.7% when the benchmark was first introduced a little over a year ago. Even without tools, Qwen 3.8 Max and its 3.7 variant show competitive scores against rivals like Opus 4.8 and Fable 5, especially in software engineering benchmarks (FrontierSWE and SWE-Pro), where it leads its direct open-source rival by a large margin.

Alibaba's commitment to releasing the model weights soon is a critical aspect, aligning with the principles of open science and open-source AI. This move is expected to empower a broader community of researchers and developers, fostering innovation and competition. The availability of powerful, cost-effective, and open-weight models like Qwen 3.8 Max signals a golden age for AI development, particularly for those with more modest computing resources. The video also features a sponsored segment for Lambda.ai, highlighting their GPU cloud services for training and inference of AI models, emphasizing the practicality of leveraging such compute resources for AI research and development.

DevOps vs SRE: delayed coverage of the dumbest war

Charity Majors · 2016-06-30 · 15 min read

TLDR: The DevOps vs SRE debate is largely a false war driven by context-blindness — Google's SRE practices are genuinely impressive but optimized specifically for Google's scale and constraints, making them largely irrelevant or even counterproductive for most companies. The real insight is that operational excellence is a shared responsibility across everyone in an org, and the "best" engineering culture is always defined by fit to context, not by imitating whatever a prestigious large company does.

Operational Best Practices #serverless

Charity Majors · 2016-05-31 · 15 min read

TLDR: Outsourcing infrastructure doesn't mean outsourcing responsibility — you still own reliability, debuggability, and failure behavior even when you don't run the systems. The critical shift is understanding your core differentiators and investing operational depth there, while accepting that opaque third-party services demand more upfront thinking about observability and resilience, not less. Blaming a vendor in a postmortem is irrelevant; you chose them, so their failures are your failures.

Simon Willison on Technical Blogging

Simon Willison · 2026-08-06 · 2 min read

Simon Willison argues that the biggest barrier to technical blogging is perfectionism, and the fix is deliberately publishing before you feel ready, because your perceived flaws are invisible to readers. For a senior engineer, this reframes blogging from a high-stakes performance into a low-friction habit, which matters because consistent public writing compounds over time into reputation, influence, and clearer thinking in ways that staying silent never does.