Beating GPT-5.6 Sol on retrieval with 100x cheaper open models
Beating GPT-5.6 Sol on retrieval with 100x cheaper open models — from Hacker News front page
New large language models ship every week — frontier models, open-weight releases, fine-tuned variants, and benchmark improvements. Keeping up requires more than a headline feed; builders need context on capability shifts, pricing changes, and API availability.
This cluster covers new model announcements, open-weight releases, benchmark claims, and model update signals from frontier labs and open-source communities. Each signal carries a link to the primary source for hands-on evaluation.
Use this page to maintain a model watchlist: identify capability jumps, track open-weight availability, and decide which models deserve a test run against your current workloads.
Beating GPT-5.6 Sol on retrieval with 100x cheaper open models — from Hacker News front page
Reddit is enlisting AI to help moderate new subreddits - and eventually the rest of site. The company is introducing automated moderation tools that rely on LLMs to help mods mana…
Anthropic is building a team for designing its own custom AI chips. The Claude maker said it would co-design hardware and models to help its technology run faster and more efficie…
MacPaw is building a local version of its AI assistant Eney using Liquid AI's models.
The Trump administration's framework for assessing potential cybersecurity risks posed by advanced AI reportedly has no interest in testing open models. Axios reports that not onl…
<p>I released <a href="https://llm.datasette.io/en/stable/changelog.html#v0-32">LLM 0.32</a> this morning, the most significant new version of LLM since the initial launch of the…
<p><strong>Release:</strong> <a href="https://github.com/simonw/llm-anthropic/releases/tag/0.26">llm-anthropic 0.26</a></p> <p>Includes new features enabled by <a href="https://si…
A new SaferAI report finds Z.ai's open-weight GLM-5.2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could o…
Tool-Integrated Reasoning (TIR) enables LLMs to solve complex tasks through iterative tool interactions. However, existing reinforcement learning methods often rely on trajectory-…
Large language models can solve substantially harder reasoning problems with more inference-time compute. The term "test-time scaling," however, now covers diverse inference algor…
Optimizing compilers miss profitable transformations when their enabling semantics are absent from the analyzed program representation. We ask whether large language models (LLMs)…
We introduce Video-DeepResearch (Video-DR), extending multimodal agents from static images to continuous video streams, a setting that demands dense spatiotemporal grounding coupl…
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories…
Human input reaches language models by typing or speaking, and each channel leaves a distinct signature: orthographic noise for keyboards; for voice, disfluency from conventional…
Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction and generation. We prove unconditional sep…