view article Article NeoMME: an efficient Multimodal-native and Multilingual Encoder Hcompany • 14 days ago • 103
view article Article Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI nico-martin, Xenova • 16 days ago • 71
view article Article How Hugging Face Inference Endpoints, Jobs, and Buckets Power Search on Papers with Code nielsr • 27 days ago • 31
view article Article Wire It, Run It, Deploy It: AI Workflows in Gradio ysharma, abidlabs • 23 days ago • 48
view article Article Prefill and Decode for Concurrent Requests - Optimizing LLM Performance tngtech • Apr 16, 2025 • 97