ollama/ollama: v0.33.1
What's Changed MLX: Qwen3.8 Flash Next support cmake: make external compat patches idempotent MLX and llama.cpp update mlxrunner: add structured output support mlxrunner: avoid Metal GPU timeouts when loading models…
Daily edition · AI Tools
TILens turns technical updates into a focused daily brief: official releases, trusted reporting, and practitioner analysis, deduplicated and organized by topic.
What's Changed MLX: Qwen3.8 Flash Next support cmake: make external compat patches idempotent MLX and llama.cpp update mlxrunner: add structured output support mlxrunner: avoid Metal GPU timeouts when loading models…
Pin cython < 3.3.0 for the Windows Triton wheel build (#194914) Summary Build Triton Windows Wheel (3.15, xpu) fails in setup.py bdist_wheel with exit status 3221225477 (0xC0000005, access violation) — a hard crash, not…
Build context was missing the new cmake common utility.
MLX: Qwen3.8 Flash Next support review comments
[Testcase Refactoring] Add hw_classification in test_shape_inference.…
[Testcase Refactoring] Add hw_classification in test/fx/test_source_m…
[Testcase Refactoring] Add hw_classification in test/fx/test_subgraph…
v0.28.0 Highlights This release features 584 commits from 270 contributors (76 new)! Kimi-K3 performance push: a major optimization effort for Kimi-K3 across the stack — Decode Context Parallel (DCP) support (#50484),…
MLflow 3.15.2 is a patch release that includes several major features and improvements. Features: [Evaluation] Support immutable evaluation dataset versions (#24845, @danielseong1) [Evaluation] Add scorer_ensemble…
Revert "[fix]: derive is_gpu()/device_need_guard() from DeviceInterfa…
Description drafted with an AI assistant, quoted below; I've reviewed the change. Replaces the c10::llvm bit helpers with their C++20 <bit> equivalents in c10 and ATen. CachingHostAllocator.h and…
[inductor] lite mode: run reinplacing before decomposing the function…
Summary: Previously we temporarily disabled expandable segments for CUDA Graph due to some issues. Now, since the issues were fixed, we re-enable it Test Plan: baseline: f1125979746 QPS: 28.5k peak memory: 97.08%…
This PR is auto-generated nightly by this action. Update the pinned vllm hash. Pull Request resolved: #194672 Approved by: https://github.com/pytorchbot
This PR is auto-generated nightly by this action. Update the pinned vllm hash. Pull Request resolved: #194672 Approved by: https://github.com/pytorchbot
[fx] Clear fake_mode on a view's base in GraphPickler (#194771) (#194…
Revert "Move c10/util math helper headers into torch/headeronly (#194…
What's Changed Claude Desktop Developers can now easily configure Claude Desktop to seamlessly work with Ollama as a third-party gateway provider. Improved caching Fixed a hang where agent clients that cancel long…
Summary Move MathConstants.h, copysign.h, Load.h, BFloat16-math.h, and irange.h from c10/util into torch/headeronly/util under torch::headeronly. Leave the old c10/util paths as one-line forwarding includes, matching…
[AOTInductor] Pluggable model-container lifecycle observer interface …
Showing 1 day · 38 items available