AI PILLED

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

Read
#983
0%Merged PRs by AI agents

0 / 413 merged PRs · 90 days

AI agents
0
User accounts
400
Other bots
13

Agents

No AI-agent authors in this read.