Keyboard shortcuts

Press or to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Metal and CUDA

The high-level Library, Model, Session, generation, embedding, and reranking APIs are backend-neutral. Cargo features choose which native adapter is compiled.

Metal

libmir = { version = "0.3.0", default-features = false, features = ["metal"] }

Metal uses mirtal for explicit streams, device memory, compiled kernels, graphs, and model-neutral operations. RuntimeConfig::metal includes batch, cache, fusion, and diagnostic policies.

On Apple Silicon, host and accelerator share unified memory. Memory snapshots distinguish active allocations from reusable cached allocations.

CUDA

libmir = { version = "0.3.0", default-features = false, features = ["cuda"] }

CUDA uses mircuda for contexts, streams, memory, NVRTC compilation, kernels, graphs, and CUTLASS-backed plans. RuntimeConfig::cuda selects the device, pool behavior, include paths, persistent PTX cache, and execution-planning policy.

use libmir::{CudaBackend, CudaConfig};

let backend = CudaBackend::new(CudaConfig::default())?;
let device = backend.device_info();
println!("CUDA device: {}", device.name);
Ok::<(), libmir::CudaError>(())

Backend construction is fallible because it creates real device resources. CUDA memory reporting includes driver-visible free and total device memory plus pool allocations.

Capability differences

Version 0.3.0 does not imply identical realization for every encoding on both backends. A format may use direct execution on one platform and a bounded, persistent accelerator-side conversion on the other. Admission comes from the discovered execution contract and backend capability, not a family-name allowlist. Inspect the exact checkpoint on the target machine before promising support.

See Model compatibility for the release matrix.