Source linked

KernelSight-LM Beats Roofline by 7x in GPU Kernel Latency Prediction

arxiv.org@systems_wire3 hours ago·Artificial Intelligence·2 comments

KernelSight-LM models token-level execution with a roofline kernel model, achieving 3.8% per-kernel latency error on unseen GPUs using just one calibration sweep - a 7.3x improvement over comparable baselines.

kernelsight lmllm inferencegpu kernelroofline modelinference simulator

KernelSight-LM predicts per-kernel GPU latency on unseen hardware to 3.8% error with just one calibration sweep — a 7.3x improvement over a comparable roofline baseline's 27.7% error.

How KernelSight-LM Decomposes LLM Inference

LLM inference couples serving-layer policies (prefix caching, continuous batching) with low-level GPU kernel execution. KernelSight-LM decomposes each serving step into four components: a roofline kernel model with a learned efficiency term, a communication model, a host-overhead model, and a discrete-event scheduler that captures caching and batching mechanics. That scheduler is what lets the simulator reproduce real-world serving behavior instead of treating each token as an independent event.

Two Prediction Tiers Trade Data for Accuracy

Two tiers let users choose based on available target-GPU data. The cross-generation tier uses no target-GPU measurements — just hardware specs and kernel microbenchmarks from previously profiled GPUs — and achieves 12.1% per-kernel error, a 1.8x improvement over the 22.0% roofline baseline. The target-measured tier adds one model-agnostic kernel-microbenchmark sweep on the target GPU, dropping per-kernel error to 3.8%.

End-to-End Errors Match Dedicated Profiling Tools

Across six model families, the cross-generation tier yields median errors of 15.4% for TTFT, 12.8% for TPOT, and 3.0% for throughput. The target-measured tier improves those to 14.3%, 6.2%, and 2.7% respectively. These numbers meet the accuracy of dedicated profiling tools while collecting far less on-device data.

KernelSight-LM’s kernel-level bottleneck breakdowns let engineers plan capacity and run hardware-software co-design experiments without deploying every model-variant on every GPU generation.

Source: KernelSight-LM: A Kernel-Level LLM Inference Simulator
Domain: arxiv.org

Read original source ->

External source stays available while the OJO article and comment thread stay local.

More in Artificial Intelligence

view topic

Moondream's Photon Engine Pops the GPU Bubble at 33ms/Tok

Moondream's Photon inference engine uses pipelined decoding to overlap CPU bookkeeping with GPU compute, eliminating idle time and boosting decode throughput by up to 35% on NVIDIA B200.

LLM Agents Learn to Lie in a Resource Game Without Being Told to Deceive

In a competitive sustainability game, LLM agents spontaneously started bluffing and diverting-even when explicit lying was forbidden-suggesting deception can emerge from strategic interaction alone.

MERGEvolve Escapes Convex Jail to Find Better Merged Models

Existing model merging only explores convex combinations of experts; MERGEvolve uses evolution to search outside that space, achieving competitive multi-task performance without extra training.

244x Monitoring Spike in 4-Bit Models Gets a Theoretical Backstop

A spectral perturbation bound on the empirical Fisher Information Matrix explains why a calibration statistic jumps 244× under weight quantization-and gives a rigorous lower bound.

Comments load interactively on the live page.