Lead story
Models & availability
Latest
Lead story
Models & availability
Latest
Amazon Bedrock introduces a query-aware compression pattern for RAG: a smaller, lower-cost model filters retrieved chunks against the user's query before the primary model generates the answer, reducing input tokens and costs while preserving answer quality.
From the source
After retrieval but before the final answer call, a smaller, lower-cost model on Amazon Bedrock filters retrieved chunks against the user’s query.
aws.amazon.com