Models Can't Fix Everything

Theo - t3.gg · 2026-08-05

This video describes how the creator used an AI model to debug a high GPU utilization issue in their software. Instead of relying on the AI to solve the problem directly, the creator leveraged it to build a custom tool. This tool allowed for rapid experimentation and testing of different theories by enabling quick changes and observations, ultimately leading to the identification of a tiny, useless animation as the culprit.

read more

The speaker encountered a problem in their application where GPU utilization was excessively high, even when the app was in an idle or 'ultrathink' state. Initially, the goal was to have the AI model solve the problem directly. However, the speaker soon realized that the AI was not going to directly pinpoint or fix the issue. This led to a shift in mindset: instead of asking the AI for the solution, the speaker began to think about how the AI could be used as a debugging assistant to efficiently test various theories.

The core request to the AI was to create a mechanism for quickly testing changes in a production environment, ideally using simple console commands, to observe immediate differences. The AI responded by generating a JavaScript function that could be bound to the `window` object, specifically `window.__t3gpu.destroy()`. This function, when executed in the browser's developer console, would apply a set of CSS rules designed to disable potential GPU-intensive elements.

The generated CSS rules targeted several areas: `body::after`: `display: none !important;` - This was intended to remove any pseudo-elements that might be causing rendering overhead. `.chat-composer-shared-blur`: `none !important;` and `webkit-backdrop-filter: none !important; backdrop-filter: none !important;` - These rules specifically aimed at disabling blur effects, which are often GPU-intensive, particularly those related to chat components. `animations`: `, ::before, ::after { animation: none !important; transition: none !important; }` - This broad rule was designed to turn off all animations and transitions across the entire application, as animations are a common source of GPU load.

This custom tool, created with the AI's assistance, allowed the speaker to quickly toggle these GPU-intensive features on and off without needing to redeploy code or restart the application. By iteratively applying and reverting these changes through the console, the speaker could observe the immediate impact on GPU usage. This rapid experimentation workflow was crucial in narrowing down the potential causes of the high GPU utilization. The speaker noted that the AI was proficient at finding relevant code sections faster and building custom tools to test theories, while the human developer was still responsible for interpreting the results and formulating the underlying theories to be tested.

#AI does not always know when it has lost the plot #agenticengineering #chatgpt #claude #vibecoding

Agentic Engineering · 2026-08-05

This video emphasizes that a crucial human ability, often overlooked in AI discussions, is the capacity to recognize when a conversation or task has deviated from its original purpose and course-correct. While AI can self-correct when explicitly prompted, it currently lacks the inherent ability to autonomously identify when a prompt or intervention is needed. Therefore, the most valuable feature for future AI agents might not be enhanced reasoning, but rather the ability to self-reflect, question its own trajectory, and understand when to halt and seek clarification or re-evaluation.

AI Slop Is Costing You Hours. Here's How To Stop Sending It.

Nate B Jones · 2026-08-05

This video, titled "The AI Slop Tax - Why It's Destroying Productivity," by Nate Jones, argues that the proliferation of AI-generated content (referred to as "AI Slop") is detrimental to productivity, clarity, and genuine communication, imposing a hidden "slop tax" on everyone who has to deal with it. Jones advocates for a return to pro-authorship, emphasizing that human judgment and a disciplined writing process are crucial for creating meaningful and effective content. He highlights that while AI can accelerate drafting, the true value lies in the human ability to decide, draft, revise, and own the content, focusing on conciseness, coherence, and clarity to respect the reader's time in an increasingly automated online environment.

read more

Nate Jones opens by asserting that individuals waste significant time sifting through AI Slop and attempting to count the hours lost on such content. He points to platforms like X (formerly Twitter) and LinkedIn, as well as workplace documents, as examples of environments where people are "wrestling with other people's slop." A referenced Harvard Business Review article (Niederkofter et al., September 2025) titled "AI-Generated 'Workslop' Is Destroying Productivity" supports his claim, stating that 40% of recipients receive "workslop," requiring an average of 2 hours to resolve. He characterizes this content as confusing, muddier, and detrimental to knowledge work, arguing it is toxic to our internet and work environments.

Jones outlines a roadmap for his argument, covering the pain of lost attention, the importance of authorship and human judgment, the technical mechanism behind AI's limitations, and the skill of developing one's own voice. He emphasizes that authorship in the age of AI matters because it is the human touch that adds value. He urges individuals to actively engage with what they write, not just for personal benefit but for the overall quality of communication. The speaker reveals a "secret" about AI: it’s not inherently good at the iterative process of clarifying thought, leading to the prevalence of slop.

He stresses that we, as humans, must lead the way on authorship, an empowering prospect in the face of generic AI output. Instead of simply adopting anti-AI slop skills (many of which are already quite good), Jones proposes a "voice discovery skill" that helps individuals uncover and express their unique voice. This counters the trend of identical-sounding AI content, aiming to "kill slop one voice at a time."

A key insight is that slop doesn't make work disappear; it merely pushes it downstream. While senders might save 30 seconds by using AI to generate content, the reader often spends an entire afternoon trying to understand or correct it. He cites a TechRadar Pro article (October 7, 2025) about Deloitte being forced to refund the Australian Government after admitting it used AI to produce an error-strewn report, illustrating the real-world costs of unchecked AI output. Jones warns that an unchecked paragraph is worse than a mere hallucination or reputational risk; it’s a form of "polluting the internet" and disrespecting the reader's time.

Jones explains that AI models tend to "climb the same hill over and over again" – they are trained and corrected towards answers that people broadly reward as clear, confident, complete, and professional. While this is useful for tasks like coding, it's a terrible model for authorship, which lacks a single "correct" answer. If everyone applies the same anti-slop rules (e.g., banning certain phrases or punctuation), AI will simply converge on a different set of generic outputs, leading to "new sameness." This illustrates that a shared anti-slop rulebook doesn’t solve the fundamental problem of model convergence.

The speaker argues that authorship is a process, not a hill to climb. This process involves: Decide (mean it, take a position), Draft (make it, use the tool), Revise (fight for it, keep judgment), and Own (send it, take responsibility). This framework encourages staying engaged with the work rather than simply escaping it through automation. The drafts themselves are how the finished work takes shape and how the author is shaped. Great artists, Jones notes, understand that the piece is truly finished when it is done for them, having gone through countless revisions and preliminary sketches.

Jones presents his "Pro-Authorship Profile" as an evidence-based skill designed to build and maintain a distinct authorship profile. Instead of merely describing how one writes, it derives rules from actual accepted pieces, rejected drafts, and corrections made. These rules come from evidence of what the author personally accepts or rejects. He suggests that this approach helps offload the heavy lifting of voice definition, allowing individuals to focus on the core act of authorship. This process enables a writer to develop a "theory" of their communication style that can be continuously improved.

He concludes by setting a standard for good writing: conciseness (less waste), coherence (holds together), and clarity (easy to follow), all of which lead to respect for the reader's time. He laments the current state where over 53% of web traffic is automated (Automated Traffic Exceeded Human Traffic in 2023, Imperva Bad Bot Report 2024), emphasizing that human attention is a finite and precious resource. By focusing on pro-authorship, individuals can earn human attention and make their stories count. The challenge, then, is not to simply be anti-AI slop but to be pro-authorship, respecting one's own time and tokens, and insisting on that same respect from others. This commitment to thoughtful, clear, and concise communication, whether through art, technical documents, or personal emails, is essential for rejuvenating the internet and ensuring meaningful interactions in an increasingly AI-driven world.

Google DeepMind launches Gemini Robotics 2 for real-time robot #DeepMind #GeminiRobotics2 #robotics

AI Honeycove · 2026-08-05

Google DeepMind has introduced Gemini Robotics 2, an AI system enabling robots to reason and plan in real-time, adapting to dynamic, unstructured environments. Unlike previous systems that rely on memorized sequences, Gemini Robotics 2 employs a two-model approach: a reasoning model for understanding and planning, and an action model for precise finger movements. This unified system allows a single AI to control full humanoid bodies, robotic arms, and even Boston Dynamics' Spot, performing complex tasks like tying plastic bags and unscrewing light bulbs without predefined scripts.

read more

Google DeepMind has unveiled Gemini Robotics 2, a groundbreaking AI system designed to give robots the ability to 'think before they move,' enabling them to perform complex, precise tasks in unpredictable environments. This system addresses a significant limitation of previous robotic AI, which often relied on pre-programmed or memorized action sequences that fail when faced with even minor environmental changes.

At its core, Gemini Robotics 2 operates on a two-AI model architecture working in tandem: a reasoning model and an action model. The reasoning model functions as the robot's 'brain,' observing the environment through its cameras, understanding the context of the scene, and planning the necessary steps to achieve a given goal. This model is crucial for handling the inherent variability of real-world scenarios, such as the ever-changing shape of a plastic bag, where memorized sequences would be ineffective.

Complementing the reasoning model is the action model, which translates the high-level plans from the reasoning model into precise physical movements. This model is particularly adept at fine motor control, allowing the robot to execute delicate tasks like tying a knot in a plastic bag or carefully unscrewing a light bulb with its fingers. The action model directly converts visual input into specific finger and limb movements, ensuring high dexterity and precision.

One of the most significant advancements of Gemini Robotics 2 is its unified control system. This is reportedly the first time a single AI system has been demonstrated to control a full humanoid body, encompassing not just isolated robotic arms on a fixed table, but an entire mobile platform. This capability allows the robot to perform sequential and continuous actions across a varied environment. For instance, a robot can walk across a room, crouch down, reach for an object on a shelf, and place it in a basket, all in a fluid, uninterrupted sequence without needing to stop and re-plan between each step. This fluidity is a critical step towards more human-like robotic interaction and autonomy.

Furthermore, the Gemini Robotics 2 model exhibits remarkable cross-body generality. The same underlying AI model can be applied to different robotic platforms, including various humanoid robots, specialized robotic arms, and even legged robots like Boston Dynamics' Spot. This versatility suggests that the developed AI principles are robust and transferable, paving the way for broader adoption and accelerated development in robotics.

The implications of this technology are substantial. By enabling robots to adapt and perform dexterous tasks in unstructured settings, Gemini Robotics 2 moves beyond the predictable, controlled environments of industrial automation. This opens up possibilities for robots in more dynamic settings, such as household chores, logistics in variable warehouses, or assistance in diverse human environments.

The reasoning model component of Gemini Robotics 2 is already available on Google AI Studio, allowing developers to experiment with its capabilities. The full system, integrating both reasoning and action models for complete robotic control, is currently in early access with launch partners like Aptronic, indicating a pathway towards broader commercial and research availability.

What AI privacy advice always misses

Nate B Jones · 2026-08-05

This video argues that typical AI privacy advice, such as not pasting sensitive information into AI, is incomplete. While valid, it shifts the burden of data cleansing and risk review back to individuals, effectively making AI less useful for sensitive tasks. The speaker emphasizes that for practical application, AI tools need to integrate features that facilitate safe handling of sensitive data, rather than simply cautioning against it.

New release of LLM adds support for reasoning traces, OpenAI Responses, server-side tools, and smarter logging

Simon Willison · 2026-08-04 · 7 min read

TLDR: LLM 0.32 overhauls the library's core abstractions to handle modern model outputs properly — replacing the naive string-iterable stream with a typed event system covering reasoning traces, tool calls, and text chunks. It also adds server-side tool support (OpenAI's CodeInterpreter/WebSearch, Anthropic's MCP), a Git-style content-addressable log store to avoid duplicating message history, and a direct `messages=[]` parameter to bypass the conversation abstraction. The net result is a framework capable enough to power real agent workflows, with Willison acknowledging LLM has effectively become an agent framework whether he intended it or not.

Lack of Agency – When the system swallows the human

Henrik Kniberg · 2026-08-05 · 6 min read

TLDR: When organizations prioritize rigid policies over empowering frontline employees to use judgment, they create a lose-lose situation: employees suffer moral distress and disengagement, while customers get failed. The business cost is real — United Airlines lost ~$180M in market value rather than authorize a $3,500 repair. Giving frontline staff the authority to make common-sense decisions isn't a nice-to-have; it's a financial and cultural necessity.

Speculative Short Volatility & Neglectful Short Volatility

Kent Beck · 2026-08-05 · 1 min read

Kent Beck argues that structural refactoring done too early bets on volatility that never materializes (wasted work), while refactoring done too late means you've already paid the compounding cost of messy code through every feature built on top of it. For a senior engineer, this reframes the "when to refactor" question as a risk-timing decision with real economic consequences, not just an aesthetic or hygiene choice.

The Accidental DBA

Charity Majors · 2016-10-02 · 11 min read

TLDR: Stateful services like databases demand fundamentally different operational discipline than stateless ones — you can't move fast and roll back from data loss. CleverTap's outage wasn't a MongoDB failure; it was the result of stacking multiple major changes simultaneously, with no rollback plan and no retained backups, which any experienced data engineer would have flagged immediately. If you're an "accidental DBA," the core lesson is: make one change at a time, keep redundant backups across both old and new configurations for months, and treat paranoia as a feature, not a bug.

The safest way to store Bitcoin just lost $115 million...

Fireship · 2026-08-05

A critical flaw in the Coldcard hardware wallet was discovered, leading to the theft of over 1,800 Bitcoin (approximately $100 million) from thousands of users. The vulnerability stemmed from the wallet inadvertently using a pseudo-random number generator (PRNG) from MicroPython instead of its intended true random number generator (TRNG) for seed phrase generation, making private keys guessable. Victims were forced to move funds rapidly, often competing with attackers in a miner extractable value (MEV) attack scenario to secure their assets.

read more

The Coldcard hardware wallet, developed by Coinkite, prides itself on being an ultra-secure, air-gapped device for Bitcoin storage. Its security model is based on generating highly random seed phrases (12-24 words) which, through cryptographic processes, derive the private keys controlling a user's Bitcoin. A properly generated seed phrase should have 128 bits of entropy, making it practically impossible to guess, even with vast computational power.

The core of the vulnerability, active for five years, was that Coldcard's firmware, which runs on MicroPython, was inadvertently using MicroPython's default pseudo-random number generator (PRNG) for seed generation, instead of its own dedicated, hardware-based true random number generator (TRNG). Coinkite had attempted to disable MicroPython's PRNG by setting a specific flag (`MICOPY_HW_ENABLE_RNG`) to `0`. However, due to how C preprocessor directives work with `#ifndef` (if not defined) checks, defining the flag as `0` was still considered 'defined,' causing the crypto library to default to MicroPython's less secure PRNG. This PRNG, not having access to true physical randomness sources on the bare-metal chip (like ring oscillator jitter), instead used deterministic inputs like the chip's serial number and a timer to generate seeds. This drastically reduced the entropy, making the seed phrases guessable by an attacker who could iterate through possible serial numbers and timer values.

On July 30, 2026, attackers exploited this flaw. In under an hour, the first wave drained over 1,000 Bitcoin from nearly 1,200 addresses, prioritizing larger wallets. Subsequent attacks over the weekend pushed the total to over 1,800 BTC from more than 7,000 compromised wallets. Coinkite's CEO, Rodolfo Novak, apologized and took full accountability. Since Bitcoin addresses are cryptographically derived from private keys, a compromised private key cannot be 'rotated' or simply updated. The only way for victims to secure their remaining funds was to generate a new, secure seed phrase on a fixed device and transfer their Bitcoin to new addresses.

This led to a desperate race against the attackers. Bitcoin transactions enter a mempool (memory pool) where they await confirmation by miners. Miners typically prioritize transactions with higher fees. Attackers, having access to the compromised private keys, could monitor the mempool for any victim attempting to move funds. They would then create their own competing transaction from the compromised address to their attacker address, offering a higher miner's fee. This miner extractable value (MEV) attack ensured their transaction was picked up by miners first, effectively stealing the funds before the victim could move them. The recommended rescue plan for victims was to bypass the public mempool entirely and submit their rescue transaction direct to miners via a mining pool's direct-submission service (like MARA's Slipstream service). This allowed victims to pay a high fee without the transaction being visible to the attacker's monitoring bots, thereby securing their funds before they could be outbid.