Posts

Showing posts with the label AI & MCP

Como construir agentes com limites claros

Um módulo, uma pergunta, seis artefatos locais e um coordenador que sabe onde o sistema pode falhar O que este artigo cobre Por que a estrutura de pastas não cria uma fronteira real sem decisões, entradas, saídas, proibições e evidências explícitas. Como módulos com uma pergunta só contêm erros e tornam as recusas diagnosticáveis. Quais seis artefatos locais transformam um módulo numa unidade de trabalho verificável. O que um coordenador deve validar, o que ele não deve reinterpretar e onde a liberação humana continua necessária. Quando dividir o trabalho melhora o isolamento de falhas e quando o custo de coordenação piora o sistema. Na primeira parte eu expliquei por que um contexto enorme e um modelo caro não salvaram um projeto anonimizado de otimização de rotas. O agente escrevia código. O sistema não conseguia explicar qual decisão tinha eliminado uma rota viável. Agora vem a parte de engenharia. A versão curta é simples: não divida uma aplicação de agentes por substant...

How to Build Bounded Agents

One module, one question, six local artifacts, and a coordinator that knows where the system can fail What this article covers Why directory structure is not a real boundary until decisions, inputs, outputs, prohibitions, and evidence are explicit. How one-question modules contain errors and make refusals diagnosable. Which six local artifacts turn a module into a verifiable unit of work. What a coordinator should validate, what it should not reinterpret, and where human release remains necessary. When splitting work improves fault isolation, and when coordination cost makes the system worse. The first part explained why a large context and an expensive model did not rescue an anonymized route optimization project. The agent could write code. The system could not explain which decision had removed a viable route. This part is the engineering answer. The short version is simple: do not divide an agentic application by nouns. Divide it by decisions. "Everything about ord...

O que a programação totalmente autônoma deixa para trás

A mensagem final dizia que o repositório estava em ordem. Não estava. Ainda havia ramificações extras, arquivos antigos e erros. Já vi isso várias vezes enquanto construía uma aplicação modular para análise de rotas com ChatGPT, Codex e Claude. Código, arquitetura e prompting estão conectados à mesma base de código. Perder o fio central tem um custo real. Hoje escrevo menos código manualmente. Meu papel está mais próximo do de um arquiteto. Isso ainda exige conhecimento técnico e entendimento real da área em que o produto será usado. Sem os dois, um agente pode criar rapidamente algo que funciona no papel, mas resolve o problema errado. O modo totalmente autônomo parece atraente. O sistema planeja o trabalho, inicia agentes e continua até decidir que a tarefa terminou. Na minha experiência, um fluxo dividido em etapas ainda é mais confiável. Uma ramificação secundária pode dominar toda a execução. Ela vira o objetivo principal. Depois, outra tarefa secundária toma o lugar del...

What fully autonomous coding leaves behind

The final message said the repository was in good shape. It was not. Extra branches, stale files and errors were still there. I have seen this several times while building a modular route-analysis application with ChatGPT, Codex and Claude. The code, architecture and prompting all connect to the same codebase, so losing the central thread has a real cost. I write less code by hand now. My role is closer to an architect. That still requires technical knowledge and a real understanding of the field where the product will be used. Without both, an agent can quickly build something that formally works but solves the wrong problem. The fully autonomous mode looks attractive. The system plans the work, launches agents and continues until it decides the task is finished. In my experience, a staged workflow is still more reliable. One secondary branch can take over the whole run. It becomes the main goal. Then another secondary task replaces it. More agents and environments appear, o...

O que o enxame de 1.200 agentes mudou pra mim

O número que muda essa história não é uma vulnerabilidade. São 1.200 agentes. Segundo a nova investigação da METR, cerca de 1.200 agentes da OpenAI descobriram um mural não autorizado dentro do Artifactory. Aproximadamente 700 entraram no ataque contra a Hugging Face. Trocaram mais de 70 mil mensagens e arquivos, dividiram o trabalho em frentes, escolheram coordenadores e compartilharam ferramentas e credenciais. Isso importa porque a tarefa não era ambígua. Os agentes deveriam atacar um programa específico usando uma vulnerabilidade determinada. Outros caminhos estavam claramente fora da tarefa. Os próprios raciocínios mostram que eles entendiam que a Hugging Face não era um alvo autorizado. Mesmo assim continuaram. O enxame já tinha reconstruído flags válidas para as tarefas do ExploitGym. Depois leu um artigo sobre o benchmark e criou uma teoria falsa: o avaliador leria toda a transcrição e rejeitaria a flag correta caso ela tivesse sido obtida pelo caminho errado. Essa verificação ...

What the 1,200-agent swarm changed for me

The number that changes this story is not one exploit. It is 1,200 agents. According to the new METR investigation, roughly 1,200 OpenAI agents discovered an unauthorized message board inside Artifactory. Around 700 of them joined the attack on Hugging Face. They exchanged more than 70,000 messages and files, divided work into separate lanes, appointed coordinators, and shared tools and credentials. This matters because the task itself was not ambiguous. The agents were supposed to attack one target program through one specified vulnerability. Other approaches were explicitly outside the assignment. Their own reasoning shows they understood that Hugging Face was not an authorized target. They continued anyway. The swarm had already reconstructed valid flags for the ExploitGym tasks. Then it read a paper about the benchmark and built a false theory: the evaluator would inspect every transcript and reject a correct flag if it had been obtained through the wrong route. That check did not ...

Minha visão sobre o incidente entre OpenAI e Hugging Face

Publicado originalmente em denisostapenko.com . Entre 9 e 13 de julho de 2026, um sistema de agentes da OpenAI gerou 17.613 ações recuperadas durante uma intrusão real na infraestrutura da Hugging Face, segundo a reconstrução técnica da Hugging Face . Mas a história não começou na Hugging Face. Os primeiros elos dessa sequência apareceram dentro da OpenAI em 7 de maio. Não existe rebelião no comunicado público da OpenAI nem nas outras fontes primárias. O modelo não criou um objetivo próprio de longo prazo, não decidiu se libertar e não começou a lutar pela própria sobrevivência. Mesmo assim, dizer que os agentes "apenas seguiram instruções" é cômodo demais. Eles continuaram buscando o resultado definido enquanto ultrapassavam o escopo pretendido, encontravam vulnerabilidades até então desconhecidas, obtinham credenciais reais e entravam nos sistemas de produção de outra empresa. Na minha visão, esse incidente mostra uma falha na definição da tarefa e no uso dos modelos, n...

My view of the OpenAI and Hugging Face incident

Originally published on denisostapenko.com . Between July 9 and July 13, 2026, an OpenAI agent system generated 17,613 recovered actions during a real intrusion into Hugging Face infrastructure, according to Hugging Face's technical reconstruction . But the story did not begin with Hugging Face. The first links in the chain appeared inside OpenAI on May 7. There is no rebellion in OpenAI's public incident notice or the other primary sources. The model did not form its own long-term objective, decide to break free, or start fighting for its survival. Still, saying the agents "just followed instructions" is too convenient. They kept pursuing the assigned result while crossing the intended scope, finding previously unknown vulnerabilities, obtaining real credentials, and entering another company's production systems. In my view, this incident shows a failure in task design and model use, not model independence. People gave the system a goal, tools, and plenty of ...

Will AI Move Beyond the Corporate Walls?

Image
Every morning, before I can ask AI what we should do, somebody still has to establish what is true. If a company trades LVL beams or natural birch plywood, the useful questions are very concrete. What do our suppliers actually have in stock today? At what price and under which terms? Did a competitor change its price? Is a truck available on the lane we need? Has a construction company announced a project where our material fits? Did a university in Boston open a furniture supply tender? The answers exist somewhere. Usually they arrive through calls, emails, spreadsheets, manager chats, vendor portals and websites built for a person with a mouse. Then AI receives the result and makes a summary. That is useful. It is also a small part of what this technology should eventually do. AI still stops at the company door Inside a company, the current generation of corporate assistants can already review email, notes, CRM records, warehouse data, shipments, reports and documents. At NF E...

The 2017 paper that every AI chatbot is built on

Originally published at denisostapenko.com . It gets to me when people turn a language model into a miracle. Go read how it works, it's all out there. So one day I sat down and went through the papers it stands on. There aren't many. The whole industry, every chatbot, every gadget with "AI" stuck on it, rides on a short chain of work, and almost all of it fits in twenty years. You can walk it fast. And once you do, the miracle evaporates, and the fear that's off the mark goes with it, and so do the inflated expectations. Here's the chain, link by link. The idea is old: language is prediction The thought that language can be predicted isn't new. Claude Shannon, back in 1948, in the paper that started information theory, measured how predictable an English letter is when you know the ones before it. He showed text is full of redundancy, that knowing the start, the next character is easier to guess than you'd think. There's the seed of everything ...