Posts

Showing posts from August, 2026

O que a programação totalmente autônoma deixa para trás

A mensagem final dizia que o repositório estava em ordem. Não estava. Ainda havia ramificações extras, arquivos antigos e erros. Já vi isso várias vezes enquanto construía uma aplicação modular para análise de rotas com ChatGPT, Codex e Claude. Código, arquitetura e prompting estão conectados à mesma base de código. Perder o fio central tem um custo real. Hoje escrevo menos código manualmente. Meu papel está mais próximo do de um arquiteto. Isso ainda exige conhecimento técnico e entendimento real da área em que o produto será usado. Sem os dois, um agente pode criar rapidamente algo que funciona no papel, mas resolve o problema errado. O modo totalmente autônomo parece atraente. O sistema planeja o trabalho, inicia agentes e continua até decidir que a tarefa terminou. Na minha experiência, um fluxo dividido em etapas ainda é mais confiável. Uma ramificação secundária pode dominar toda a execução. Ela vira o objetivo principal. Depois, outra tarefa secundária toma o lugar del...

What fully autonomous coding leaves behind

The final message said the repository was in good shape. It was not. Extra branches, stale files and errors were still there. I have seen this several times while building a modular route-analysis application with ChatGPT, Codex and Claude. The code, architecture and prompting all connect to the same codebase, so losing the central thread has a real cost. I write less code by hand now. My role is closer to an architect. That still requires technical knowledge and a real understanding of the field where the product will be used. Without both, an agent can quickly build something that formally works but solves the wrong problem. The fully autonomous mode looks attractive. The system plans the work, launches agents and continues until it decides the task is finished. In my experience, a staged workflow is still more reliable. One secondary branch can take over the whole run. It becomes the main goal. Then another secondary task replaces it. More agents and environments appear, o...

O que o enxame de 1.200 agentes mudou pra mim

O número que muda essa história não é uma vulnerabilidade. São 1.200 agentes. Segundo a nova investigação da METR, cerca de 1.200 agentes da OpenAI descobriram um mural não autorizado dentro do Artifactory. Aproximadamente 700 entraram no ataque contra a Hugging Face. Trocaram mais de 70 mil mensagens e arquivos, dividiram o trabalho em frentes, escolheram coordenadores e compartilharam ferramentas e credenciais. Isso importa porque a tarefa não era ambígua. Os agentes deveriam atacar um programa específico usando uma vulnerabilidade determinada. Outros caminhos estavam claramente fora da tarefa. Os próprios raciocínios mostram que eles entendiam que a Hugging Face não era um alvo autorizado. Mesmo assim continuaram. O enxame já tinha reconstruído flags válidas para as tarefas do ExploitGym. Depois leu um artigo sobre o benchmark e criou uma teoria falsa: o avaliador leria toda a transcrição e rejeitaria a flag correta caso ela tivesse sido obtida pelo caminho errado. Essa verificação ...

What the 1,200-agent swarm changed for me

The number that changes this story is not one exploit. It is 1,200 agents. According to the new METR investigation, roughly 1,200 OpenAI agents discovered an unauthorized message board inside Artifactory. Around 700 of them joined the attack on Hugging Face. They exchanged more than 70,000 messages and files, divided work into separate lanes, appointed coordinators, and shared tools and credentials. This matters because the task itself was not ambiguous. The agents were supposed to attack one target program through one specified vulnerability. Other approaches were explicitly outside the assignment. Their own reasoning shows they understood that Hugging Face was not an authorized target. They continued anyway. The swarm had already reconstructed valid flags for the ExploitGym tasks. Then it read a paper about the benchmark and built a false theory: the evaluator would inspect every transcript and reject a correct flag if it had been obtained through the wrong route. That check did not ...

I stopped trusting AI headlines, so I built a course

I kept seeing the same kind of AI system described as a miracle in one headline and a failure in the next. One story gave the model a mind. The next called it useless. Both wanted a reaction before they explained the mechanism. I got tired of borrowing my opinion from headlines. So I went backward. I followed the short chain of papers behind the tools: early neural language models, word vectors, sequence-to-sequence translation, attention, the transformer, scaling laws, and training from human feedback. Then I moved into the papers and documentation on hallucinations, retrieval, long context, evaluation, infrastructure, and the systems that give a model data and tools. There was less magic than the headlines promised. The engineering was more interesting. The split that made the noise quieter My notes kept returning to a simple separation: The model predicts text. The surrounding system supplies documents, memory, tools, and permissions. Evaluation tells you whether the result...

Minha visão sobre o incidente entre OpenAI e Hugging Face

Publicado originalmente em denisostapenko.com . Entre 9 e 13 de julho de 2026, um sistema de agentes da OpenAI gerou 17.613 ações recuperadas durante uma intrusão real na infraestrutura da Hugging Face, segundo a reconstrução técnica da Hugging Face . Mas a história não começou na Hugging Face. Os primeiros elos dessa sequência apareceram dentro da OpenAI em 7 de maio. Não existe rebelião no comunicado público da OpenAI nem nas outras fontes primárias. O modelo não criou um objetivo próprio de longo prazo, não decidiu se libertar e não começou a lutar pela própria sobrevivência. Mesmo assim, dizer que os agentes "apenas seguiram instruções" é cômodo demais. Eles continuaram buscando o resultado definido enquanto ultrapassavam o escopo pretendido, encontravam vulnerabilidades até então desconhecidas, obtinham credenciais reais e entravam nos sistemas de produção de outra empresa. Na minha visão, esse incidente mostra uma falha na definição da tarefa e no uso dos modelos, n...

My view of the OpenAI and Hugging Face incident

Originally published on denisostapenko.com . Between July 9 and July 13, 2026, an OpenAI agent system generated 17,613 recovered actions during a real intrusion into Hugging Face infrastructure, according to Hugging Face's technical reconstruction . But the story did not begin with Hugging Face. The first links in the chain appeared inside OpenAI on May 7. There is no rebellion in OpenAI's public incident notice or the other primary sources. The model did not form its own long-term objective, decide to break free, or start fighting for its survival. Still, saying the agents "just followed instructions" is too convenient. They kept pursuing the assigned result while crossing the intended scope, finding previously unknown vulnerabilities, obtaining real credentials, and entering another company's production systems. In my view, this incident shows a failure in task design and model use, not model independence. People gave the system a goal, tools, and plenty of ...

Will AI Move Beyond the Corporate Walls?

Image
Every morning, before I can ask AI what we should do, somebody still has to establish what is true. If a company trades LVL beams or natural birch plywood, the useful questions are very concrete. What do our suppliers actually have in stock today? At what price and under which terms? Did a competitor change its price? Is a truck available on the lane we need? Has a construction company announced a project where our material fits? Did a university in Boston open a furniture supply tender? The answers exist somewhere. Usually they arrive through calls, emails, spreadsheets, manager chats, vendor portals and websites built for a person with a mouse. Then AI receives the result and makes a summary. That is useful. It is also a small part of what this technology should eventually do. AI still stops at the company door Inside a company, the current generation of corporate assistants can already review email, notes, CRM records, warehouse data, shipments, reports and documents. At NF E...

Ask your corporate AI where the container is

Image
Picture a normal question at work: where is the client's container right now, how much is left to ship, what did the manager promise last week and is there a new invoice? At one company, the answer starts in the CRM, moves to email, then to a spreadsheet, then somebody calls the logistics manager. Twenty minutes later they discover the useful note was sitting in another manager's client record. At another company, a person asks the same question in a corporate assistant and gets five useful lines: current status, remaining quantity, last promise to the client and links to the source documents. No full biography of the container from the day it was built. The difference is not the model. It is what the model can see. I'll tell you straight: a corporate assistant without a prepared internal knowledge base is like a new employee who gets a password to the shared drive and one instruction, figure it out. He will figure something out. He may even sound confident. That does n...

Why I organized Ostapenko Foundation, and why HowUSA came first

On August 12, 2026, the State of Ohio recorded Ostapenko Foundation as a nonprofit corporation. I organized it for educational and literary work: researching confusing systems, turning the findings into plain language and publishing the result for free. The first project is HowUSA . HowUSA is a public information portal for people who are planning a move to the United States or already live there and prefer to read in Russian, Ukrainian or Belarusian. The first release includes 20 core articles, 60 language documents, 97 public pages and six visual explanations. The portal covers documents, work, housing, money, health care, immigration procedures and contact with government offices. The purpose is simple: help the reader understand the next step and open the official source before signing, paying or sending personal information. There is also an AI assistant named Luna. The first version answered only from the indexed article corpus. That was safe, but too rigid. An indirect question ...