Following a recipe is a skill, but it isn’t cooking. A line cook can execute a Beef Wellington perfectly every time, but if the oven temperature fluctuates or the meat is poor quality, they are stuck. A chef, however, understands the chemistry of heat and protein. They can pivot because they understand the “why” behind the process. Most of what we call “AI for science” right now is just a very fast line cook. It is pattern matching on a massive scale, which is fine for summarizing papers or predicting a protein structure based on known folds, but it is useless for actually discovering something that isn’t already hiding in the training set.

The gap between data retrieval and actual reasoning is the real wall. As noted in MIT Tech Review, the push is now toward agents that can reason—meaning they can hypothesize, test, and fail. But here is the rub: reasoning requires a level of autonomy that makes corporate legal teams sweat. You cannot have a “reasoning” agent that is simultaneously constrained by a thousand layers of RLHF designed to prevent it from saying something mildly offensive to a shareholder. If a model is taught to be cautious above all else, it will never suggest the radical, counter-intuitive hypothesis that actually leads to a discovery.

This brings us to the “censorship-industrial complex.” We have seen this movie before with the early days of image generators (remember the “safety” filters that blocked everything including the word ‘breast’ in a medical context?). Now we are applying that same corporate sanitization to scientific discovery. When you wrap a model in so many safety rails that it cannot even discuss certain chemical precursors or biological pathways for fear of “dual-use” risks, you aren’t making the world safer. You are just making the tool useless for the people who actually know how to use it. Why are we pretending that a model managed by a PR department is the right tool for a high-stakes lab? (I suspect most of these “safety” committees have never stepped foot in a wet lab).

The current obsession with “alignment” in science models is a mistake. Alignment is for chatbots that talk to the general public, not for agents running simulations on materials science or quantum chemistry. The goal of a science agent should be accuracy and raw capability, not politeness. If a model is too “safe” to suggest a volatile but necessary reaction, it is not a scientific tool; it is a toy. We are essentially trying to build a telescope that refuses to look at certain parts of the sky because those regions might be “controversial.” It is an absurd approach to discovery that prioritizes corporate optics over actual progress.

Then there is the friction of the hardware. Running a true reasoning agent—one that iterates, self-corrects, and runs internal simulations—isn’t a cheap API call. It is an iterative loop that eats tokens like candy and requires massive VRAM overhead to maintain the state of a complex experiment. The latency alone on these complex chains is enough to make you miss the era of simple Python scripts. If the goal is to move beyond data-munching, we have to accept that these models will be slower and significantly more expensive to run than the flashy chat interfaces we are used to.

Most “AI for science” initiatives are just marketing exercises for compute clusters.

The industry is at a breaking point. The tension between the need for raw, uncensored reasoning and the corporate need for safety is unsustainable. By Q1 2027, we will see a major “jailbroken” or natively uncensored science model released by an open-source collective that renders the current corporate “safe” versions obsolete. Until then, we will keep pretending that a model that apologizes for its inability to help is actually helping science.