The article discusses the impact of concurrent request processing on LLM latency, throughput, and GPU utilization, focusing on the prefill and decode phases.
Introducing HELMET: Holistically Evaluating Long-context Language Models
This blog introduces HELMET, a comprehensive benchmark for evaluating long-context language models (LCLMs) that addresses limitations of existing evaluations by providing diverse, controllable, and…
Cohere is now a supported Inference Provider on Hugging Face Hub, allowing serverless inference for a range of Cohere and Cohere Labs models including Command variants, Aya Expanse, and Aya Vision.
The article explains 17 features that make Gradio distinct, positioning it as a framework for building interactive AI/ML applications with built-in APIs, server-side rendering, queue management,…
Hugging Face and Protect AI have been partnering for six months to improve model security using Guardian's scanning technology, resulting in over 4 million models scanned and new threat detection…
OpenAI o3 and o4-mini combine state-of-the-art reasoning with full tool capabilities including web browsing, Python, image and file analysis, image generation, canvas, automations, file search, and…
OpenAI introduced GPT-4.1, a new family of models with improvements in coding, instruction following, and long-context understanding, along with their first nano model.
LG AI Research presented a paper at AAAI-25 introducing ImagePiece, a content-aware re-tokenization technique for Vision Transformers that re-tokenizes non-semantic visual tokens into semantic units,…
LG AI Research presented a paper titled 'VarDrop: Enhancing Training Efficiency by Reducing Variate Redundancy in Periodic Time Series Forecasting' at the AAAI-25 main track, introducing a method to…