Amazon announces two new APIs for SageMaker Feature Store: BatchWriteRecord for batch writing up to 25 records across multiple feature groups, and ListRecords for enumerating record identifiers…
Decathlon, a major sporting goods retailer, selected Chronos-2 as a core component of its forecasting stack after evaluating multiple time series foundation models.
AWS announced the SchedulingConfig parameter for the CreateInferenceComponent API in Amazon SageMaker AI Inference Components, giving customers fine-grained control over IC copy placement across…
Amazon announces an integration between Amazon Quick and fal via the Model Context Protocol (MCP), enabling agentic creative workflows for media teams.
Amazon Bedrock announces the availability of OpenAI GPT-5.6 models (Terra and Luna) in India with India geographic cross-Region inference, allowing customers to process inference requests and data…
Amazon Bedrock now supports OpenAI GPT-5.6 models (Terra and Luna) in India with India geographic cross-Region inference, enabling data residency within India.
Deepgram announces Enhanced Metrics and Prometheus/OpenTelemetry support for its speech-to-text and text-to-speech models deployed on Amazon SageMaker AI, providing billing and usage transparency and…
Amazon Web Services, NVIDIA, and Heidi Health announce a solution that uses NVIDIA CUDA Multi-Process Service and NVIDIA Triton Inference Server on Amazon EC2 GPU instances to reduce GPU…
Amazon Bedrock AgentCore Evaluations is a new evaluation service that uses OpenTelemetry to score any agent framework (LangGraph, LlamaIndex, OpenAI Agents SDK, Google ADK, Claude Agent SDK, Strands…
Natera built an intelligent appointment scheduling voice agent using Amazon Bedrock AgentCore, achieving 100% tool-calling accuracy and sub-7-second latency at less than $0.01 per call.
Amazon announces script mode in SageMaker Python SDK v3, which allows users to bring their own model without building Docker images, using new unified classes ModelTrainer and ModelBuilder, and a…
The post is a guide on advanced data preparation strategies for supervised fine-tuning, covering learning curve analysis, data subset selection, augmentation, and mixing, with references to Amazon…
A blog post describing how to connect Amazon Bedrock AgentCore agents in one AWS account to knowledge bases in another account using Amazon Redshift Serverless, with two orchestration variants: a…
Amazon OpenSearch Service announced MCP Apps, a capability that extends the Model Context Protocol to return interactive visualizations alongside text responses in AI assistant chats, enabling…
Amazon announces Amazon Quick Desktop, a native desktop application that extends Amazon Quick with AI-assisted reporting capabilities, including local file access, background processing, and a…
Amazon announces new Ray capabilities on SageMaker HyperPod, integrating Ray with HyperPod's purpose-built infrastructure for foundation model training and serving.
This post presents a customizable, smart-caching cloud-based solution for knowledge management using AWS services, including Amazon Bedrock Knowledge Bases, Amazon Titan Text Embeddings, and an…
Amazon announces the Agentic Resource Discovery (ARD) open specification for federated agent discovery across environments, and the AWS Agent Registry, a centralized catalog for agents, MCP servers,…
Amazon Web Services published a blog post providing a walkthrough for building a restaurant telephony AI host using Amazon Connect, Amazon Lex V2, Amazon Connect Agentic Voice, Amazon Connect…
This blog post demonstrates an AI-powered metadata correction and harmonization workflow built on AWS, using Amazon Bedrock for LLM-powered schema alignment and correction recommendations.
OpenAI and Thailand's MHESI launch an eight-week accelerator to help 10 health, wellness, and education startups turn AI prototypes into trusted products.
OpenAI published a study involving over 1,000 students that examines the impact of ChatGPT and critical-thinking training on originality and student performance in a real-world university assignment.
OpenAI is expanding ChatGPT for Teachers to 55 U.S. school systems, providing secure AI tools, training, and support to over 100,000 educators and staff.
This post is a customer story about how loveholidays uses OpenAI Codex to enable software development across the business, allowing teams to quickly turn ideas into products.
OpenAI's CFO discusses the compounding effects of advances across hardware, compute, models, and products to deliver more useful intelligence at greater scale and lower cost.
OpenAI announced Jalapeño, a custom inference chip that offers faster, more power-efficient inference with higher throughput and lower latency for modern models.
OpenAI introduces the Admin plugin for ChatGPT Work and Codex, enabling workspace usage analysis, member and permission management, limit adjustments, and handling of admin requests.
OpenAI enforced its safety policy by banning accounts originating from Russia that were leveraging AI to promote a fabricated think tank and a sovereignty index that favored Russia and disparaged the…
August 2026 brought more control over how GitHub Copilot reasons, which models you use, how teams share specialized agents, and when you ask for a code review.
GitHub Copilot weekly releases on August 24 include new features such as team sessions in Slack and Teams, a generally available Customize tab that brings MCP servers, plugins, skills, and canvases…
GitHub announces upcoming changes to Copilot policies and billing: billing changes for Copilot Business/Enterprise starting September/October 2026, convergence of Copilot Chat experiences into a…
GitHub Copilot code review now supports reviewing pull requests authored by bots (including Copilot cloud agent) and very large pull requests, and adds the ability to specify a resolution reason…
Enterprise-managed settings now support opting individual plugin marketplaces into automatic updates by setting autoUpdate: true, reducing manual maintenance for organization customizations.
GitHub announced the general availability of the Customize tab in the GitHub Copilot app, which unifies MCP servers, plugins, skills, and canvases in a single interface for discovering and managing…
The Open ASR Leaderboard adds two new evaluation sets, Monsoon en-IN and Monsoon hi-IN, for Hindi and Indian English, to address bias and improve representation of Global South languages and diverse…
Sentence Transformers v6.0 introduces a new model type, MultiVectorEncoder, for ColBERT-style late interaction retrieval, along with a complete training approach.
IBM announced the Granite 4.2 family of reasoning LLMs in three sizes (3B, 8B, 30B), pre-trained from scratch on about 15T tokens, with a context window extended to 512K tokens, supervised…
IBM announces the release of two new Granite Speech 5.0 Turbo CTC speech recognition models: a 470M-parameter Apache 2.0 licensed model and a CC-BY-NC-SA-4.0 licensed model, both offering high…
Introduces Quantization-Aware Healing (QAH), a method that distills from the original pre-compression model to recover capabilities lost during structural compression and quantization.
Gradio introduces gr.Workflow, a new feature that allows users to build AI pipelines as a graph of typed nodes, which serves as a drag-and-drop interface, a REST API, and can be deployed to Hugging…
This blog post provides a general overview of generative AI for business, covering use cases, benefits, and challenges. It does not announce any new product, model, or feature.
Cohere discusses how forward-deployed engineers (FDEs) can build customer capability to avoid vendor lock-in, emphasizing co-building and knowledge transfer rather than creating operational…
Cohere announces Parse, a cost-effective vision language model for enterprise document parsing, available via API and Model Vault, with pricing at $1.50 per 1,000 pages.
Cohere announces findings from a commissioned IDC study on sovereign AI adoption, highlighting definitional gaps, top concerns (data leakage and compliance), and the need for a holistic strategy.
NVIDIA announced NVHBM, a custom high-bandwidth memory technology that integrates the memory controller into the HBM stack, delivering up to 30% greater memory bandwidth, 15% lower power consumption,…
NVIDIA announced NVLink Fusion, an infrastructure solution that connects custom XPUs to NVIDIA's NVLink scale-up domain and AI factory ecosystem, enabling faster time to market and reduced risk for…
NVIDIA announced that its Groq 3 LPX inference accelerator is in full production, delivering 3,400 output tokens per second on Gemma 4 31B for 100,000-token long-context use cases, 4x faster than…
NVIDIA announces measured performance data for Vera Rubin NVL72 systems, showing up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads and up to 35x lower cost per million…
Perplexity announced support for the new model perplexity/deepseek-v4-flash-0731 in its Agent API and Router API, describing it as a fast, efficient open reasoning model with a 1M-token context…
Perplexity's Agent API and Router API now support the perplexity/nemotron-3.5-lightning-30b-a3b model, an open-weight reasoning model, with pricing details.
Anthropic announces 10,000 free or discounted Claude subscriptions for scientists via a new team plan, expands AI for Science program with up to $50,000 in credits per project, and notes restrictions…
Anthropic is previewing the Model Hardware Standard (MHS), a shared specification enabling AI agents to safely operate physical devices like microscopes and robotic arms, reducing integration time…
Anthropic announces a $5 million grant program to fund independent research into how AI impacts users' wellbeing, providing direct funding, model access, and technical support to grantees building…
Antigravity announces several updates to Teamwork, a multi-agent orchestration framework, including new patterns like Iterative Coding, Distributed Coding, Long Proof, Self-Verification, and Document…
Google Antigravity announces Generative UI and visual tooling that allows agents to create, render, and iterate on rich visual components, dynamic dashboards, interactive artifacts, and design…
The Antigravity 2.0 coding assistant now includes a VCS side panel and an embedded terminal, allowing users to perform Git operations, view diffs, stage and commit changes, and run command line tools…
Google DeepMind announces Gemini Omni 1.1 Flash, a new suite of creative controls and generative video capabilities including scene extension, first and last frame interpolation, 360p draft mode, 4K…
Google DeepMind announces the world's first double-blind evaluation of a proprietary frontier AI model, using cryptographic environments to prevent benchmark contamination.
Google DeepMind announces Gemini 3.5 Transcribe, a new speech-to-text model for real-time transcription, available via the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform.
Google Research introduces the planetary prediction engine (PPE), an experimental research capability under the Google Earth AI initiative that autonomously executes the full geospatial modeling…
Google Research announces GlucoFM, a lightweight self-supervised foundation model for continuous glucose monitoring that uses a dual-stream design to separate slower glycemic trends from short-term…
AgentHands is a research prototype that uses LLMs to generate synchronized hand gestures for AI agents in XR, enabling spatially grounded conversations.
Meta announced MetaRoCE, a new RDMA transport protocol designed for AI workloads on commodity Ethernet, and is releasing its specification, reference software implementation, and compliance test…
Meta announces MTIA 300, its first training chip with built-in NIC chiplets and communication-offloading engines, designed to optimize communication for training recommendation models, achieving 3.9x…
Microsoft publishes a comparison of managed PostgreSQL (Azure Database for PostgreSQL, Azure HorizonDB) versus self-hosted PostgreSQL, discussing trade-offs in control, operational tax, security, and…
Microsoft Foundry blog post discusses four strategies to optimize agent costs: selecting the right model per request, caching, prompt and agent optimization, and observability.
Grok Bot now integrates with X, allowing users to connect their X account, get free X API credits, and perform actions like searching posts, reading timelines, and checking mentions.
LG AI Research introduces ReSQL, a self-improving Text-to-SQL framework that uses retrieval-augmented error reasoning to correct queries, demonstrating error reductions on benchmarks.
Mistral AI and HUMAIN announced a strategic collaboration to advance sovereign AI in Saudi Arabia and the Middle East, focusing on AI infrastructure, model development, and local deployment in areas…
Tencent releases Hy4 preview, a 770B-parameter model with 49B active parameters and a 1M-token context window, open-sourced and available on Tencent Cloud and OpenRouter.
An analysis comparing GLM-5.3 and its distilled sibling GLM-5.3 Flash on the DeepSWE benchmark, finding that the Flash variant retains 94% of solved tasks at one-seventeenth the cost but loses…