State-of-the-art LLM compression, built for production inference with vLLM
Reading the pull requests.
0 / 203 merged PRs · 90 days
No AI-agent authors in this read.