ollama/ollama: v0.33.1
What's Changed MLX: Qwen3.8 Flash Next support cmake: make external compat patches idempotent MLX and llama.cpp update mlxrunner: add structured output support mlxrunner: avoid Metal GPU timeouts when loading models…
Topic - Edition
What's Changed MLX: Qwen3.8 Flash Next support cmake: make external compat patches idempotent MLX and llama.cpp update mlxrunner: add structured output support mlxrunner: avoid Metal GPU timeouts when loading models…
Qwen3.8-Flash-Next Another open weights model from Qwen. This one is "a multimodal MoE model that also serves as an early preview of the architecture used in Qwen4". It's pretty big: 125B tokens, but only 6B active…
3.4.0 (2026-08-25) Features api: Add obfuscation field to ChatCompletionChunk (#3690) (c7d8e1d) api: add project residency configuration and cost quantity units (#3726) (bc4f8ef) Bug Fixes api: encode Realtime call…
Health & Bioscience
Pin cython < 3.3.0 for the Windows Triton wheel build (#194914) Summary Build Triton Windows Wheel (3.15, xpu) fails in setup.py bdist_wheel with exit status 3221225477 (0xC0000005, access violation) — a hard crash, not…
Build context was missing the new cmake common utility.
1.1.0 (2026-08-26) Full Changelog: v1.0.0...v1.1.0 Features api: add updates thinking display mode (beta) (eb4a73f) api: add missing anthropic-beta values (dbebd15) api: add support for Organization API endpoints…
MLX: Qwen3.8 Flash Next support review comments
Release v5.16.0 New Model additions Qwen4-Exp Qwen4-Exp builds on Qwen3.5's hybrid text and multimodal architecture with three key components: GatedResidual (GR), Qwen Sparse Attention (QSA), and Per-Layer Embedding…
Release v5.16.1 This is a special release as we include GLM! (and a few small fixes) GLM-5.3-Flash GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active…
[Testcase Refactoring] Add hw_classification in test_shape_inference.…
[Testcase Refactoring] Add hw_classification in test/fx/test_source_m…
[Testcase Refactoring] Add hw_classification in test/fx/test_subgraph…
v0.28.0 Highlights This release features 584 commits from 270 contributors (76 new)! Kimi-K3 performance push: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484),…
ChatGPT for Teachers is expanding to 55 U.S. school systems, bringing secure AI tools, training, and support to over 100,000 more educators and staff.
OpenAI’s new report explores how students and educators use ChatGPT to make learning more continuous, with support that extends beyond the classroom.
MLflow 3.15.2 is a patch release that includes several major features and improvements. Features: [Evaluation] Support immutable evaluation dataset versions (#24845, @danielseong1) [Evaluation] Add scorer_ensemble…
The fact that AI wrote 1M LOC and then refined it over the course of the next couple of months to produce a reliable piece of software that is currently running on millions of developer machines is absolutely mind…
Revert "[fix]: derive is_gpu()/device_need_guard() from DeviceInterfa…
Description drafted with an AI assistant, quoted below; I've reviewed the change. Replaces the c10::llvm bit helpers with their C++20 <bit> equivalents in c10 and ATen. CachingHostAllocator.h and…
100 items available