Hugging Face researchers introduce LAVE, an LLM-assisted evaluation metric for zero-shot VQA on the Docmatix dataset, questioning whether fine-tuning is necessary for accurate evaluation.
Meta released Llama 3.1 models in three sizes (8B, 70B, 405B) with base and instruct variants, featuring 128K context length, multilingual support, tool calling, and a more permissive license.
Hugging Face published a blog post detailing how to run Mistral 7B with Core ML on a Mac, using new features from WWDC 24 such as Swift Tensor and Stateful Buffers, and achieving less than 4GB memory…
OpenAI announces it is testing SearchGPT, a temporary prototype of new search features that provide fast and timely answers along with clear and relevant sources.
Replicate announces API support for running Meta's Llama 3.1 405B model, including usage examples in JavaScript, Python, and cURL, as well as details on the model's parameters and context window.
Improving Model Safety Behavior with Rule-Based Rewards
OpenAI developed a new method called Rule-Based Rewards (RBRs) to align models for safe behavior without needing extensive human data collection.