Skip to content
TILens What matters today in tech v0.1.0
Theme

Daily edition · AI

The daily ledger

TILens turns technical updates into a focused daily brief: official releases, trusted reporting, and practitioner analysis, deduplicated and organized by topic.

31 Aug 2026 edition
AI GitHub AI Tools

ggml-org/llama.cpp: b10729

metal : add fa-vec tunings for M1 Ultra (#28088) metal : add fa-vec tunings for M1 Ultra metal : move M1 Ultra tunings after M1 Max section metal : remove duplicate blank line Website: https://llama.app Attestations:…

Source: llama.cpp Releases github-actions[bot]
AI GitHub AI Tools

ggml-org/llama.cpp: b10728

CUDA: XOR swizzle flash attn K,V smem fp16 tiles (#25635) CUDA: XOR swizzle flash attn K,V smem fp16 tiles Signed-off-by: ynankani ynankani@nvidia.com Fix use 64bit generic pointer instead of 32bit shared pointer…

Source: llama.cpp Releases github-actions[bot]
AI GitHub AI Tools

ggml-org/llama.cpp: b10727

metal : add concat support for quantized types (#28116) Assisted-by: pi:llama.cpp/DeepSeek-V4-Flash-0731 Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/44252653 macOS/iOS:…

Source: llama.cpp Releases github-actions[bot]
AI GitHub AI Tools

ggml-org/llama.cpp: b10726

AVX2: Speed up large batch size prompt processing of IQ models (#27402) Batched gemm for grid IQ quants Style updates and a bit more performance Clean up comments Move code around Vectorize IQ panel decode, lower…

Source: llama.cpp Releases github-actions[bot]
AI GitHub AI Tools

ggml-org/llama.cpp: b10724

kv-cache : optimize restoring non-contiguous cells (#27991) kv cache : batch state restore scatter reads per contiguous run When restoring state into non-contiguous destination cells (e.g. a prompt-cache snapshot into a…

Source: llama.cpp Releases github-actions[bot]
AI GitHub AI Tools

ggml-org/llama.cpp: b10723

opencl: tune the quant paths for Intel Xe-LP GPUs to improve its TG and PP performance (#26438) opencl: Q4_K/Q5_K mul_mv N_DST 4->8 on Intel for 2x activation reuse opencl: Q4_K mul_mm 8x8 tile fot Intel opencl: Q5_K…

Source: llama.cpp Releases github-actions[bot]
AI GitHub AI Tools

ggml-org/llama.cpp: b10721

webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation (#28045) webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_backend_tensor_get() implementation…

Source: llama.cpp Releases github-actions[bot]
AI GitHub AI Tools

ggml-org/llama.cpp: b10720

ROCm: add radix TOP_K for long rows (#27466) ROCm: add radix TOP_K for long rows Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/44178618 macOS/iOS: macOS Apple Silicon…

Source: llama.cpp Releases github-actions[bot]
AI GitHub AI Tools

ggml-org/llama.cpp: b10719

metal : add fa-vec tunings for M1 (#28078) Website: https://llama.app Attestations: https://github.com/ggml-org/llama.cpp/attestations/44172949 macOS/iOS: macOS Apple Silicon (arm64) macOS Apple Silicon (arm64, KleidiAI…

Source: llama.cpp Releases github-actions[bot]

Showing 1 day · 424 items available