OpenAI introduced Trusted Access for Cyber, a trust-based framework that expands access to frontier cyber capabilities while strengthening safeguards against misuse.
OpenAI announced OpenAI Frontier, an enterprise platform for building, deploying, and managing AI agents with features like shared context, onboarding, permissions, and governance.
OpenAI announces GPT-5.3-Codex, described as the most capable agentic coding model to date, combining the coding performance of GPT-5.2-Codex with the reasoning and knowledge capabilities of GPT-5.2.
OpenAI introduced GPT-5.3-Codex, a Codex-native agent that combines frontier coding performance with general reasoning for long-horizon, real-world technical work.
OpenAI describes how to embed the Codex agent using the Codex App Server, a bidirectional JSON-RPC API for streaming progress, tool use, approvals, and diffs.
OpenAI describes the philosophy behind the Sora feed, emphasizing creativity, connection, safety, personalized recommendations, parental controls, and guardrails.
OpenAI and Snowflake have entered a $200M partnership to integrate frontier AI capabilities into Snowflake's enterprise data platform, enabling AI agents and insights within Snowflake.
OpenAI introduced the Codex app for macOS, a command center for AI coding and software development featuring multiple agents, parallel workflows, and long-running tasks.
ServiceNow-AI announces SyGra Studio, an interactive visual environment for synthetic data generation workflows that lets users compose flows on a canvas, preview datasets, tune prompts with inline…
Hugging Face announces Community Evals, a new feature on the Hugging Face Hub that allows benchmark datasets to host leaderboards, models to store their own eval scores, and the community to submit…
H Company announces Holo2-235B-A22B Preview, a new UI localization model achieving 78.5% on Screenspot-Pro and 79.0% on OSWorld G, available on Hugging Face as a research release.
This blog post analyzes the development of the open-source AI ecosystem in China since the DeepSeek Moment, examining the strategies of major Chinese AI organizations such as Alibaba, Tencent,…
This is the second part of a series on training efficient text-to-image models from scratch. It documents training techniques that improved convergence and representation learning for the PRX model,…
A study of near-unconstrained LLM generation reveals that different model families exhibit distinct topical preferences: GPT-OSS favors programming and math, Llama leans literary, DeepSeek often…
Together AI announces the availability of two new Rime models: Rime Arcana V3 Turbo (English–Spanish, performance) and Rime Arcana V3 (11-language switching), both supporting native code-switching…
Together AI announces the appointment of Alon Gavrielov as VP of Infrastructure Strategy, who will lead the expansion of AI factories to support growing AI-native applications.
Together Evaluations now supports proprietary models from OpenAI, Anthropic, and Google as both judges and targets, allows evaluation of Together fine-tuned models, and provides new recipes…
Open-source LLM judges fine-tuned with DPO can outperform GPT-5.2 at evaluating model outputs. Together AI trained GPT-OSS 120B on 5,400 preference pairs to beat GPT-5.2's accuracy, delivering…
Google Research introduces Natively Adaptive Interfaces (NAI), a framework for creating more accessible applications through multimodal AI tools that adapt to user needs, co-developed with the…
Google Research introduces Sequential Attention, a subset selection algorithm for making large-scale ML models more efficient by greedily selecting features, layers or blocks during training without…
Google Research announces a partnership with Included Health to launch a nationwide randomized study evaluating conversational AI in real-world virtual care, pending IRB approval.
Qwen3-Coder-Next is an open-weight language model designed for coding agents and local development, built on Qwen3-Next-80B-A3B-Base with hybrid attention and MoE.
Baidu announced ERNIE 5.0, a 2.4 trillion parameter multimodal foundation model with a unified autoregressive framework integrating text, images, video, and audio, achieving state-of-the-art results…
LG AI Research announced the graduation of two PhD students from its LG Graduate School of AI, highlighting their research on video action recognition and monocular depth estimation, and the school's…
Moonshot AI releases WorldVQA, a benchmark of 3,500 image-question pairs to evaluate factual visual world knowledge in multimodal LLMs, with a focus on head vs. tail distribution.
Tencent Hunyuan introduces CL-bench, a benchmark to evaluate whether language models can learn new knowledge from context rather than relying solely on parametric memory, covering four types of…
Z.ai launched GLM-OCR, a compact and high-performance optical character recognition model using a self-developed CogViT and GLM-0.5B encoder-decoder architecture with CLIP pre-training for robust…