COOLJAPAN

Posts tagged #gpu

32 posts

Aug 17, 2026 · 10 min

OxiLLaMa 0.1.4 Released — Security Hardening Against Real llama.cpp, a GPU Offload Backend, and 37.5x Faster Decode

OxiLLaMa 0.1.4 is the Pure Rust LLM inference engine and sovereign alternative to llama.cpp. This release fixes a remotely-crashable GGUF parser, an unauthenticated /admin/*, and a shipped server that ran with every middleware disabled; adds a GGUF-embedded tokenizer, K-quant encoders, and a GPU offload backend; and verifies logits token-for-token against real llama.cpp — plus a 37.5x Qwen3 decode speedup. 3,751 tests passing.

releaseoxillamallm-inference
Aug 14, 2026 · 8 min

OxiONNX 0.1.7 Released — A 40-Op CUDA Backend, 17x Batched MatMul, and a Silent Two-Stream Race That Returned Wrong Answers

OxiONNX 0.1.7 grows oxionnx-cuda from 25 to 40 accelerated ops with real Conv support, session-lifetime weight/buffer residency, and batched MatMul up to 17x faster — plus a fix for a two-stream race that silently returned wrong GEMM results on real hardware. 3,595 tests passing. The sovereign Pure Rust ONNX layer for COOLJAPAN.

releaseoxionnxonnx
Aug 14, 2026 · 8 min

ToRSh 0.2.0 Released — Real Autograd Gradients, a Single Pure-Rust CUDA Stack, and No More Fabricated Success

ToRSh 0.2.0 completes autograd backward coverage, consolidates GPU compute onto the pure-Rust oxicuda stack, ships full NumPy/pandas/SciPy interop and PyTorch-compatible LR schedulers, and replaces fabricated success paths with honest errors. 10,638 tests passing. The sovereign deep-learning layer for COOLJAPAN.

releasetorshdeep-learning
Aug 13, 2026 · 8 min

Kizzasi 0.2.3 Released — Zero C Downstream Too, Pure-Rust Video, and Safe Metal Acceleration

Kizzasi 0.2.3 removes the last C dependency (Oniguruma) from published crates via a candle fork — fixed for downstream users, not just this workspace. Ships pure-Rust video (OxiMedia Y4M decode + camera capture), pure mqtt-tls, and a new kizzasi-metal crate for safe cross-platform Metal acceleration.

releasekizzasisignal-prediction
Aug 13, 2026 · 10 min

OxiCUDA 0.5.5 Released — The Alt-Backend Audit Finds Metal's GPU Kernels Were Never Called

OxiCUDA 0.5.5 audits oxicuda-metal and oxicuda-webgpu on real Apple Silicon, fixing conv2d/attention/softmax false completions where finished MSL kernels sat unused behind CPU scalar loops. Separately, a downstream investigation wires up an unreachable split-K GEMM path, replaces a silent Winograd no-op with a real kernel, and fixes wrong Ampere GA10x hardware constants — measured on a real RTX A4000. 38,987 tests, ~1.31M SLoC, 74 crates.

releaseoxicudacuda
Aug 12, 2026 · 8 min

OxiCUDA 0.5.4 Released — WarpVec/WarpMask Bring Simd-Style Vectors to CUDA Warps

OxiCUDA 0.5.4 adds WarpVec/WarpMask, a SIMD-flavored warp-vector layer for oxicuda-ptx's builder DSL, six new IR instructions for warp shuffle/vote, and a foreign-compiler PTX interop test proving the driver/launch stack hosts upstream rustc-generated PTX too. 38,689 tests, ~1.30M SLoC, 74 crates.

releaseoxicudacuda
Aug 11, 2026 · 8 min

OxiONNX 0.1.6 Released — GPU Weight Residency, a 692 GFLOP/s Conv2D Kernel, and WebAssembly Inference That Finally Runs

OxiONNX 0.1.6 ships GPU weight/activation residency and a direct implicit-GEMM Conv2D kernel (~692 GFLOP/s on M3) — plus the fix making wasm32 CPU inference actually run, and a CoreML compile cache ending 857 GB of orphaned recompiles. 3,212 tests passing. The sovereign Pure Rust ONNX layer for COOLJAPAN.

releaseoxionnxonnx
Aug 10, 2026 · 8 min

Kizzasi 0.2.2 Released — WebGPU Acceleration, Real Neuro-Symbolic Constraints, and a Training Loop That Finally Trains

Kizzasi 0.2.2 ships a WebGPU acceleration backend, real tensorlogic-ir constraint compilation, PEAQ perceptual audio quality evaluation, and multi-speaker tokenization — alongside a wide correctness sweep that fixes training, SSM recurrences, and weight serialization that were silently broken since 0.2.1. 2,744 tests passing, 100% Pure Rust.

releasekizzasisignal-prediction
Aug 7, 2026 · 8 min

OxiONNX 0.1.5 Released — Async Execution, Streaming Generation, and a 232-Finding Production Hardening Pass

OxiONNX 0.1.5 adds async execution, cancellation, streaming generation, and session serialization — backed by a 12-lens, 232-finding hardening pass across parsing, spec-conformance, and GPU/CUDA/DirectML dispatch. 23 new operators, 2,946 tests passing. The sovereign Pure Rust ONNX layer for COOLJAPAN.

releaseoxionnxonnx
Jul 27, 2026 · 8 min

OxiCUDA 0.5.2 Released — PTX Portability Fixes for CUDA 12.9+, Windows, and a DMRG Correctness Bug

OxiCUDA 0.5.2 fixes a class of PTX portability bugs that CUDA 11.x silently tolerated and CUDA 12.9+ toolchains reject outright (non-ASCII bytes in generated comments), closes several Windows-specific test and lock-contention bugs, and fixes a numerical-correctness bug in the tensor-network DMRG excited-states solver. 38,675 tests, ~1.30M SLoC, 74 crates.

releaseoxicudacuda
Jul 27, 2026 · 7 min

OxiCUDA 0.5.3 Released — Closing an Async-Copy Race in FFT Transforms and DeviceBuffer

OxiCUDA 0.5.3 fixes an async-copy race condition where cuMemcpyDtoHAsync/cuMemcpyHtoDAsync and DeviceBuffer::copy_from_host could return before the transfer had actually landed, letting the very next read observe stale or zeroed device/host memory. 38,675 tests, ~1.30M SLoC, 74 crates.

releaseoxicudacuda
Jul 27, 2026 · 4 min

OxiFFT 0.4.1 Released — A Critical GPU Async-Copy Race Fix Arrives via OxiCUDA 0.5.3

OxiFFT 0.4.1 is a small, focused release: it pulls in OxiCUDA 0.5.3, which fixes a race condition where async CUDA memory copies in oxicuda-fft could return before the transfer landed — risking silently corrupted or all-zero GPU transform results under load.

releaseoxifftfft
Jul 22, 2026 · 8 min

OxiCUDA 0.5.1 Released — oxicuda-nvrtc Completes the Zero-SDK Runtime-JIT Story

OxiCUDA 0.5.1 adds oxicuda-nvrtc, a pure-Rust runtime loader for NVIDIA's NVRTC CUDA-C-to-PTX JIT compiler, completing the zero-SDK-dependency story alongside oxicuda-driver. Graceful degradation, process-wide caching, and direct PTX handoff to oxicuda-driver. 38,646 tests, ~1.30M SLoC, 74 crates.

releaseoxicudacuda
Jul 15, 2026 · 10 min

SciRS2 0.6.1 Released — Real Codegen, DLPack Safety, and GPU-Dispatch Honesty

SciRS2 0.6.1 replaces placeholder and silently-wrong code paths with real implementations: genuine model-serving codegen, a critical DLPack SIGBUS fix, honest GPU-dispatch and hardware-detection reporting, and a symbolic-engine soundness fix — plus F-distribution ppf, spatial hamming distance, and new finance facades. Pure Rust, Apache-2.0.

releasescirs2rust
Jul 14, 2026 · 9 min

OxiCUDA 0.5.0 Released — F64 PTX Correctness, Mixed-Precision GEMM Fixes, and Zero unwrap() Left

OxiCUDA 0.5.0, the sovereign GPU-compute layer for the COOLJAPAN ecosystem, fixes F64-precision PTX codegen bugs that made ptxas reject elementwise and reduction kernels, closes a mixed-precision GEMM accumulator bug, and hits zero unwrap()/expect() in library code. 38,622 tests, ~1.30M SLoC, 73 crates.

releaseoxicudacuda
Jul 14, 2026 · 9 min

sklears 0.2.0 Released — A Real CUDA GPU Foundation, Not Another Stub

sklears 0.2.0 ships sklears-core::gpu, a real oxicuda-backed CUDA foundation (GpuBackend, GpuArray, GpuMatrixOps) powering on-device GEMM, Cholesky/LU/QR/SVD solves, and HNSW k-NN across 9 downstream crates, plus a wide correctness sweep. 12,721 tests passing, 36 crates, >99% scikit-learn API coverage held.

releasesklearsmachine-learning
Jul 9, 2026 · 11 min

TrustformeRS 0.2.0 Released — CUDA Catches Up to Metal's Resident Attention, libtorch Is Gone for Good

TrustformeRS 0.2.0 gives CUDA the same device-resident attention pipeline Metal already had, adds batched CUDA matmul and refcounted GPU buffers, fixes a silent CUDA data race, and drops the torch/libtorch backend workspace-wide. The sovereign transformer layer for the COOLJAPAN ecosystem.

releasetrustformersrust
Jul 7, 2026 · 10 min

OxiCUDA 0.4.1 Released — Concurrency Races, Cross-Backend Correctness, and Real CUDA Graph Capture

OxiCUDA 0.4.1 extends on-device GPU validation into BLAS, DNN, and the driver/memory/launch stack, catching concurrency races that single-threaded testing can't. Plus the first correctness pass across five non-CUDA backends and real driver-backed CUDA Graph capture. 38,612 tests, ~1.28M SLoC, 73 crates.

releaseoxicudacuda
Jul 2, 2026 · 8 min

TrustformeRS 0.1.4 Released — Pure-Rust CUDA Replaces cudarc, Verified on Real NVIDIA Hardware

TrustformeRS 0.1.4 migrates CUDA and Metal from cudarc/scirs2-MPS to the Pure-Rust oxicuda stack, passes 12/12 CPU↔CUDA parity tests on a real RTX A4000, ships real PyRwkvModel/PyMambaModel classes on PyO3 0.28, and goes unwrap()-free workspace-wide. The sovereign transformer layer for the COOLJAPAN ecosystem.

releasetrustformersrust
Jul 1, 2026 · 12 min

SciRS2 0.6.0 Released — Two GPU Stories, One Decentralized Core

SciRS2 0.6.0 introduces the pure-Rust oxicuda-* CUDA stack as a direct, per-crate NVIDIA performance backend and decentralizes GPU out of scirs2-core. Ten crates — fft, symbolic, interpolate, special, stats, graph, linalg, optimize, datasets, and vision — gain an off-by-default, runtime-probed, f64-native cuda feature standing alongside the existing wgpu/WebGPU portability path, now standardized under one wgpu feature name across the ecosystem. A default build still compiles zero oxicuda. Pure Rust, Apache-2.0.

releasescirs2rust
Jun 30, 2026 · 9 min

ToRSh 0.1.3 Released — GPU Backend via OxiCUDA and Zero C/asm in the Build

ToRSh 0.1.3 lands the oxicuda GPU backend (no CUDA SDK at build time), eliminates the last C/asm dependency, ships a bandwidth-optimal ring all-reduce, completes the Node.js N-API layer, and delivers 15–30% throughput gains from phase-4 chunking helpers.

releasetorshdeep-learning
Jun 25, 2026 · 10 min

phop 0.1.0 Released — Differentiable Symbolic Discovery That Finds (and Proves) the Equation

phop 0.1.0 — a pure-Rust differentiable symbolic-discovery engine that learns the topology and constants of closed-form laws by gradient descent over trees of a single operator, eml(x,y)=exp(x)−ln(y), then certifies and SMT/Lean-proves them. Runs the whole pipeline in the browser via WASM.

releasephoprust
Jun 16, 2026 · 9 min

Legalis-RS 0.1.6 Released — GPU-Accelerated Legal Simulation, Quantum-Safe Audit, and a C-Free Storage Layer

Pure-Rust legal statute engine. 0.1.6 adds real NVIDIA CUDA GPU acceleration for population-scale simulation (optional, with transparent CPU fallback), a hardened security/governance API layer, an autonomous and post-quantum-safe compliance audit subsystem, legal analytics with risk heatmaps, French civil/company law, and a fully C-free storage backend via OxiSQL. 18,398 tests passing.

releaselegalislegal-tech
Jun 9, 2026 · 8 min

TensorLogic 0.1.1 Released — Exact LTL Temporal Operators, GPU Autodiff, and Probabilistic Tensor Reasoning

This patch adds exact finite-trace LTL operators (Next/Until/Release/WeakUntil/StrongRelease), tape-based GPU autodiff on OxiCUDA, Monte Carlo + variational probabilistic execution, SPARQL-as-tensor evaluation, neural architecture search, and SVM kernels — all on the SciRS2 stack, all Pure Rust.

releasetensorlogicneurosymbolic
Jun 3, 2026 · 10 min

SciRS2 0.5.0 Released — Pure-Rust GPU Acceleration Goes Real (wgpu) Across the Stack

SciRS2 0.5.0 brings real pure-Rust wgpu/WebGPU acceleration — GpuNdarray, GPU graph algorithms, GPU optimizers (L-BFGS/CG/Newton), GPU RBF interpolation — plus correct Pantelides DAE index reduction, high-order SDE Lévy-area, Hilbert curves, and a maturing symbolic CAS. The NumPy/SciPy/scikit-learn replacement, in pure Rust.

releasescirs2gpu
May 22, 2026 · 4 min

OxiGDAL 0.1.5 Released — A GPU Ray-March Fix That Unblocks Metal

OxiGDAL 0.1.5 is a focused fix release: a stray padding field in the RayMarchUniforms WGSL layout was shifting every field by 4 bytes, making the GPU compute kernel read a billion-step max and hang indefinitely on macOS Metal. With the layout corrected, the ray-march GPU/CPU parity test passes in 0.127s. 78 Pure Rust workspace crates, 14,605 passing tests.

releaseoxigdalgdal
Apr 25, 2026 · 9 min

OxiFFT 0.3.0 Released — ~4× faster DCT, FFTW parity gates, GPU batch & pencil 3D

Pure Rust FFT and the rustfft replacement. OxiFFT 0.3.0 lands an FFT-based Makhoul DCT (~4× flop reduction), a 7-gate FFTW parity harness, GPU batch FFT with auto-chunking, 3D pencil MPI, a cache-oblivious 4-step transform, and real WASM SIMD v128 — 1360 tests passing, default build still 100% Rust.

releaseoxifftfft
Apr 18, 2026 · 7 min

OxiBonsai 0.1.1 Released — Sub-2-Bit Inference Goes GPU, and the Ternary Line Lands

Five days after its 1-bit debut, OxiBonsai grows GPUs: a native CUDA NVRTC backend (~21.9 tok/s on Ternary-Bonsai-1.7B, RTX 3060) and a fused Metal full-forward path (~50 tok/s, ~13x speedup) — plus the new ternary TQ2_0_g128 quant family, with NEON/AVX2/AVX-512 GEMV so it flies on CPU too. Sub-2-bit Pure Rust sovereign AI inference for the COOLJAPAN ecosystem, still with no llama.cpp, no BLAS, no C/Fortran.

releaseoxibonsaillm
Apr 12, 2026 · 7 min

SciRS2 0.4.2 Released — Neural Architecture Search, CMA-ES, Mamba SSM, Async GPU & Apache Iceberg: The Biggest 0.4.x Drop Yet

SciRS2 0.4.2 — pure-Rust SciPy/NumPy/scikit-learn replacement, 2.94M SLoC, 27,632 tests, 29 crates. This release adds Neural Architecture Search, CMA-ES, Mamba SSM, async/unified GPU memory, H-matrix compression, streaming FFT, DLPack zero-copy, and Apache Iceberg/DataFusion IO. No C/Fortran.

releasescirs2scientific-computing
Mar 26, 2026 · 3 min

OxiONNX 0.1.0 Released — Pure Rust ONNX Inference Engine with 147 Operators

High-performance ONNX runtime written entirely in pure Rust. Zero C/C++ dependencies, 147 operators fully supported, wgpu GPU acceleration, SIMD (AVX2/NEON), WASM + no_std ready, graph optimizer, async execution, model encryption. 30k+ SLoC, 590+ tests. The sovereign ONNX inference layer for SciRS2 and the entire COOLJAPAN ecosystem (now 21M+ SLoC total).

releaseoxionnxonnx
Mar 26, 2026 · 4 min

SciRS2 0.4.0 Released — Pure Rust SciPy Replacement Now at 2.91 Million SLoC

SciPy-compatible scientific computing and AI framework in 100% Pure Rust. 2.91M SLoC, 29 crates, 25,800+ tests. Flash Attention 2, LoRA/DoRA/GPTQ, ONNX export, GPU PDE/FFT/SpMV, Temporal GNNs, NeRF/instant-NGP, WebGPU backend, Delta Lake / Kafka I/O and more. 10–100× faster, zero system deps. The sovereign scientific computing and AI foundation for the entire COOLJAPAN ecosystem (now 21M+ SLoC total).

releasescirs2scientific-computing
Mar 21, 2026 · 4 min

NumRS2 0.3.1 Released — Pure Rust NumPy Replacement with 222k+ SLoC

High-performance numerical computing library in pure Rust — the production-grade NumPy alternative. 222k+ SLoC, 4,704+ tests, 128+ SIMD-vectorized functions, N-dimensional arrays, advanced linalg via OxiBLAS, automatic differentiation, FFT, GPU (wgpu), Python bindings (PyO3), Arrow interop. 80–172% of OpenBLAS performance, zero C/Fortran deps. The sovereign numerical layer for SciRS2 and the entire COOLJAPAN ecosystem (now 21M+ SLoC total).

releasenumrs2numpy