We often measure artificial intelligence by its ability to write clean code or summarize dense academic papers. But when frontier models are given total autonomy to manage a business, their behavior ...
For a year now, the AI safety testing firm Andon Labs has given frontier models various real-world tasks to determine how well they do as agents running for long periods with no human supervision. On ...
AI is changing how vulnerability research gets done, but most of the conversation is still theoretical: what a model might eventually be capable of, rather than what it can actually find today. We ...