-
Notifications
You must be signed in to change notification settings - Fork 1.1k
New serverless pattern - bedrock-semantic-cache-s3vectors-sam #3262
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
b0a2197
43c3ec3
d9c153e
d9ce0c2
e32a392
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,151 @@ | ||
| # Serverless semantic cache for Amazon Bedrock (with Amazon S3 Vectors) | ||
|
|
||
| Return cached answers for **semantically similar** prompts - different wording still hits - so you skip the LLM call on repeats and near-repeats. Cuts Amazon Bedrock cost and latency, scales to zero, and drops in front of any model. | ||
|
|
||
| Learn more at Serverless Land Patterns: https://serverlessland.com/patterns/bedrock-semantic-cache-s3vectors-sam | ||
|
|
||
| > Important: this application uses AWS services (AWS Lambda, Amazon Bedrock, Amazon S3 Vectors, AWS Systems Manager) and there are costs associated with these services after the Free Tier usage. You are responsible for any AWS costs incurred. No warranty is implied in this example. | ||
|
|
||
| --- | ||
|
|
||
| ## TL;DR - read this first | ||
|
|
||
| - **What it is:** a Lambda in front of Amazon Bedrock that caches answers by *meaning* (not exact text). Same question asked three different ways -> one Bedrock call, two instant cache hits. | ||
| - **Best for:** FAQ / support bots, docs Q&A, high-traffic assistants - anywhere many users ask the same things in different words. | ||
| - **Not for:** answers that must be exact, fresh, or per-user (unless you add namespacing / invalidation / a verify step). | ||
| - **Cost:** a cache hit skips the expensive LLM call and pays only a tiny embedding + vector query. Break-even is a few percent hit rate, so any repetitive workload is a net win. | ||
| - **Distinct from Bedrock native features:** native Prompt Caching is *exact-prefix* only (one character breaks it); Intelligent Prompt Routing picks a cheaper model. This caches by *semantic similarity* and skips the model entirely. They complement each other. | ||
|
|
||
| ## Where it shines | ||
|
|
||
| - **Repetitive, paraphrase-heavy traffic.** Research shows ~31% of LLM queries are semantically similar to a prior one - those become instant, free hits. | ||
| - **Latency-sensitive UX.** Measured ~7x faster on a hit (~230 ms vs ~1,700 ms). | ||
| - **Cost-sensitive, high-volume assistants.** Every hit is one fewer Bedrock invocation and does **not** count against your Bedrock TPM/RPM limits (throttle relief under load). | ||
| - **Any model / provider.** The cache is model-agnostic; the on-miss call is a drop-in for Bedrock or an external model. | ||
|
|
||
| ## Where it will not shine | ||
|
|
||
| - **Unique, one-off prompts.** No repetition -> ~0 hits -> you pay a tiny per-request overhead for nothing. Skip it here. | ||
| - **Answers that must be exact or fresh.** A similar-but-not-identical prompt can return a subtly different prior answer. Mitigate with a higher threshold + TTL, or bypass the cache for such routes. | ||
| - **Per-user / personalized answers.** Namespace the cache per user, or don't cache these. | ||
| - **Semantic antonyms.** The built-in negation guard catches "not / n't", but not opposites like "cheapest" vs "most expensive". For high-stakes (financial, legal, medical), add an optional LLM equivalence-verify on borderline hits. | ||
|
|
||
| ## How it saves money regardless (the math) | ||
|
|
||
| Every request pays a tiny "cache tax": one embedding call (~$0.00002) + one vector query (fractions of a cent). You **save** on every HIT because you skip the LLM call (cents to dollars, especially with large prompts / RAG context / bigger models). | ||
|
|
||
| ``` | ||
| net savings = (hits x LLM cost skipped) - (all requests x tiny cache tax) | ||
| ``` | ||
|
|
||
| Because the tax is orders of magnitude smaller than an LLM call, **break-even is roughly a 1-5% hit rate.** Real repetitive workloads sit far above that, and savings **compound** as the cache warms. The only losing case is genuinely zero repetition. Plus: everything is pay-per-use and **scales to zero** (Lambda + S3 Vectors), so there is no idle cost. | ||
|
|
||
| --- | ||
|
|
||
| ## How it works | ||
|
|
||
| ``` | ||
| prompt --> [Lambda] --embed--> Amazon Bedrock (Titan v2 -> 1024-dim vector) | ||
| | | ||
| |--search--> Amazon S3 Vectors (cosine top-K; answer is stored in vector metadata) | ||
| | |-- HIT (sim >= threshold, fresh, current epoch, negation-parity) --> return cached answer (~230 ms, $0 LLM) | ||
| | |-- MISS --> | ||
| |--generate--> Amazon Bedrock LLM (writes the answer) | ||
| |--store-----> Amazon S3 Vectors (embedding + answer + model + created_at + epoch) | ||
| |--return | ||
| (force-invalidate epoch is stored in AWS Systems Manager Parameter Store) | ||
| ``` | ||
|
|
||
| - **Amazon Bedrock** is used two ways: **embeddings** (turn text into a meaning vector so matching is semantic) and the **LLM** (answer on a miss). | ||
| - **Amazon S3 Vectors** is the cache store *and* the similarity search - pay-per-use, no always-on cost. This is the primitive that makes a serverless semantic cache economical. | ||
| - **AWS Lambda** is stateless glue. The cache lives entirely in S3 Vectors, so it survives cold starts, redeploys, and env recycling. | ||
| - **SSM Parameter Store** holds the epoch counter for force-invalidation. | ||
|
|
||
| ### Correctness features | ||
| - **Tunable similarity threshold** (default cosine 0.85) - per-deploy and per-request. | ||
| - **Freshness TTL** - entries older than `TTL_SECONDS` are treated as a miss. | ||
| - **Force-invalidate** - bump one epoch number -> every prior entry instantly misses (no deletes/scans). For big changes that can't wait for TTL. | ||
| - **Negation-parity guard** - "is X" vs "is NOT X" embed ~identically but mean the opposite; the guard blocks that false hit. | ||
| - **top-K + iterate** - a stale duplicate near-neighbour never blocks a valid hit. | ||
|
|
||
| ## Requirements | ||
|
|
||
| - An AWS account with permissions for AWS Lambda, Amazon Bedrock, Amazon S3 Vectors, and AWS Systems Manager. | ||
| - [AWS CLI](https://docs.aws.amazon.com/cli/latest/userguide/install-cliv2.html) v2, recent enough to include the `s3vectors` commands. | ||
| - [AWS SAM CLI](https://docs.aws.amazon.com/serverless-application-model/latest/developerguide/serverless-sam-cli-install.html). | ||
| - Amazon Bedrock **model access enabled** for the embeddings model (`amazon.titan-embed-text-v2:0`) and the text model (`amazon.nova-lite-v1:0`) in your Region. | ||
| - A Region where Amazon S3 Vectors and Amazon Bedrock are available (e.g. `us-east-1`). | ||
|
|
||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. The Testing section uses ## Requirements
* ... (existing items)
* Python 3 (only needed to run the test payload helper commands below) |
||
| ## Deployment | ||
|
|
||
| S3 Vectors is not yet a CloudFormation resource, so create the vector store first (two commands), then deploy the rest with SAM. | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Amazon S3 Vectors now has CloudFormation support, so this statement is inaccurate and steers users toward manual CLI steps that could instead be codified in the template. Amazon S3 Vectors resources (AWS::S3Vectors::VectorBucket and
AWS::S3Vectors::Index) can now be defined directly in the SAM/CloudFormation
template, removing the manual setup commands. If you prefer to keep the
vector store decoupled from the stack lifecycle, the two CLI commands below
remain a valid alternative. |
||
|
|
||
| ```bash | ||
| # 1. Create the S3 Vectors bucket and a cosine index (1024 dims = Titan v2) | ||
| export VECTOR_BUCKET="semantic-cache-$(aws sts get-caller-identity --query Account --output text)" | ||
|
Comment on lines
+84
to
+85
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Deployment jumps straight to the git clone https://github.com/aws-samples/serverless-patterns
cd serverless-patterns/bedrock-semantic-cache-s3vectors-sam
# then continue with the S3 Vectors setup and sam build/deploy below |
||
| aws s3vectors create-vector-bucket --vector-bucket-name "$VECTOR_BUCKET" | ||
| aws s3vectors create-index \ | ||
| --vector-bucket-name "$VECTOR_BUCKET" \ | ||
| --index-name prompt-cache --data-type float32 --dimension 1024 --distance-metric cosine \ | ||
| --metadata-configuration 'nonFilterableMetadataKeys=prompt,response,model,created_at,epoch' | ||
|
|
||
| # 2. Build and deploy the Lambda + IAM + SSM epoch parameter | ||
| sam build | ||
| sam deploy --guided | ||
| # - VectorBucket: value of $VECTOR_BUCKET above | ||
| # - VectorIndex : prompt-cache | ||
| # - ApiKey : (optional) a secret for the x-api-key header, or leave blank for IAM-only | ||
| ``` | ||
|
|
||
| Note the `FunctionUrl` and `FunctionName` outputs. | ||
|
|
||
| ## Testing | ||
|
|
||
| The function URL uses AWS_IAM auth (SigV4). The simplest test is a direct invoke: | ||
|
|
||
| ```bash | ||
| KEY="<the ApiKey you set, or omit the header if blank>" | ||
| payload() { python3 -c "import json,sys;print(json.dumps({'headers':{'x-api-key':'$KEY'},'body':json.dumps({'prompt':sys.argv[1]})}))" "$1" > ev.json; } | ||
|
|
||
| # MISS (calls Bedrock) | ||
| payload "What is the capital of France?"; aws lambda invoke --function-name semantic-cache --cli-binary-format raw-in-base64-out --payload file://ev.json out.json; cat out.json | ||
| sleep 6 | ||
| # HIT - exact repeat (cached=true, similarity ~1.0, ~230ms) | ||
| payload "What is the capital of France?"; aws lambda invoke --function-name semantic-cache --cli-binary-format raw-in-base64-out --payload file://ev.json out.json; cat out.json | ||
| # HIT - semantic (different words) | ||
| payload "Which city is the capital of France?"; aws lambda invoke --function-name semantic-cache --cli-binary-format raw-in-base64-out --payload file://ev.json out.json; cat out.json | ||
| ``` | ||
|
|
||
| Force-invalidate (e.g., after a data/policy change): | ||
|
|
||
| ```bash | ||
| python3 -c "import json;print(json.dumps({'headers':{'x-api-key':'$KEY'},'body':json.dumps({'action':'invalidate'})}))" > ev.json | ||
| aws lambda invoke --function-name semantic-cache --cli-binary-format raw-in-base64-out --payload file://ev.json out.json; cat out.json | ||
| # -> {"invalidated": true, "epoch": N}. Every prior answer now misses (propagates within ~30s). | ||
| ``` | ||
|
|
||
| Expected: exact/semantic repeats HIT (`cached=true` with a similarity score); unrelated prompts MISS; after `invalidate`, the same prompt MISSes once, then HITs again once re-cached. | ||
|
|
||
| ## Tuning | ||
|
|
||
| | Setting | Env var / request field | Effect | | ||
| |---|---|---| | ||
| | Similarity threshold | `SIM_THRESHOLD` (deploy) or `threshold` (per request) | Higher = stricter matching, fewer but safer hits | | ||
| | Freshness | `TTL_SECONDS` | Max age of a served answer | | ||
| | Force-invalidate | `POST {"action":"invalidate"}` | Invalidate the whole cache instantly | | ||
| | Models | `EMBED_MODEL`, `LLM_MODEL` | Swap embeddings / answer model | | ||
|
|
||
| ## Cleanup | ||
|
|
||
| ```bash | ||
| sam delete | ||
| aws s3vectors delete-index --vector-bucket-name "$VECTOR_BUCKET" --index-name prompt-cache | ||
| aws s3vectors delete-vector-bucket --vector-bucket-name "$VECTOR_BUCKET" | ||
| aws ssm delete-parameter --name /semantic-cache/epoch | ||
| ``` | ||
|
|
||
| --- | ||
|
|
||
| Author: Manish S | ||
|
|
||
| Copyright 2026 Amazon.com, Inc. or its affiliates. All Rights Reserved. SPDX-License-Identifier: MIT-0 | ||
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,69 @@ | ||
| { | ||
| "title": "Serverless semantic cache for Amazon Bedrock with Amazon S3 Vectors", | ||
| "description": "Cut Amazon Bedrock cost and latency by returning cached answers for semantically-similar prompts, using AWS Lambda and Amazon S3 Vectors.", | ||
| "language": "Python", | ||
| "level": "300", | ||
| "framework": "AWS SAM", | ||
| "patternArch": { | ||
| "icon1": { "x": 20, "y": 50, "service": "lambda", "label": "AWS Lambda (semantic cache)" }, | ||
| "icon2": { "x": 55, "y": 30, "service": "bedrock", "label": "Amazon Bedrock (embeddings + LLM)" }, | ||
| "icon3": { "x": 55, "y": 70, "service": "s3", "label": "Amazon S3 Vectors (cache store)" }, | ||
| "line1": { "from": "icon1", "to": "icon2" }, | ||
| "line2": { "from": "icon1", "to": "icon3" } | ||
| }, | ||
| "introBox": { | ||
| "headline": "How it works", | ||
| "text": [ | ||
| "Large language model (LLM) calls are slow and expensive, yet a large share of production prompts are paraphrases of ones already answered. This pattern places an AWS Lambda function in front of Amazon Bedrock that returns a cached answer whenever an incoming prompt is semantically similar to a previous one - so you skip the LLM call entirely on repeats and near-repeats.", | ||
| "When a request arrives, the Lambda function embeds the prompt with an Amazon Bedrock embeddings model (Amazon Titan Text Embeddings v2, 1024 dimensions). It then queries an Amazon S3 Vectors index for the nearest stored prompt using cosine similarity. Amazon S3 Vectors returns the closest match together with its distance and metadata in a single call, and the cached answer is stored directly in that metadata - so no separate database is required.", | ||
| "On a cache HIT (cosine similarity at or above a configurable threshold, the entry still within its freshness TTL, the current cache epoch, and passing a negation-parity guard) the function returns the stored answer in milliseconds at zero LLM cost. On a MISS, the function calls the Bedrock text model to generate the answer, stores the prompt embedding plus the answer, model, timestamp and epoch back into Amazon S3 Vectors, and returns the fresh result.", | ||
| "Correctness is designed, not assumed. A tunable similarity threshold trades precision for hit rate. A freshness TTL bounds staleness. A one-call force-invalidation bumps a global epoch stored in AWS Systems Manager Parameter Store, so every previously cached answer instantly becomes a miss without deleting or scanning anything - ideal for a policy or data change that cannot wait for the TTL. A negation-parity guard prevents the classic semantic-cache trap where a prompt and its negation ('is X' versus 'is NOT X') embed almost identically but mean the opposite. The query uses top-K retrieval and iterates candidates, so a stale duplicate near-neighbour never blocks a valid hit.", | ||
| "Amazon S3 Vectors is what makes this economical and fully serverless: pay-per-use vector storage and search that scales to zero, instead of an always-on vector database. The AWS Lambda function is stateless - the cache persists entirely in Amazon S3 Vectors, so it survives cold starts, redeploys, and execution-environment recycling. Access is over an IAM-signed Lambda function URL, with an optional application-level API key. IAM permissions follow least privilege: bedrock:InvokeModel is scoped to foundation models; s3vectors actions (QueryVectors, GetVectors, ListVectors, PutVectors, GetIndex) are scoped to the specific vector bucket and index; and ssm:GetParameter/PutParameter is scoped to the single epoch parameter.", | ||
| "The embeddings model and the answer model are both configurable, and the on-miss call is a drop-in for any model provider - the caching layer itself is model-agnostic. Best fit: FAQ and support assistants, documentation Q&A, and high-traffic assistants where users ask the same things in different words. Not intended for answers that must be exact, fresh, or per-user unless combined with namespacing, invalidation, or an equivalence-verification step." | ||
| ] | ||
| }, | ||
| "gitHub": { | ||
| "template": { | ||
| "repoURL": "https://github.com/aws-samples/serverless-patterns/tree/main/bedrock-semantic-cache-s3vectors-sam", | ||
| "templateURL": "serverless-patterns/bedrock-semantic-cache-s3vectors-sam", | ||
| "projectFolder": "bedrock-semantic-cache-s3vectors-sam", | ||
| "templateFile": "template.yaml" | ||
| } | ||
| }, | ||
| "resources": { | ||
| "headline": "Additional resources", | ||
| "bullets": [ | ||
| { "text": "Amazon S3 Vectors - vector storage in Amazon S3", "link": "https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-vectors.html" }, | ||
| { "text": "Amazon Bedrock - Titan Text Embeddings", "link": "https://docs.aws.amazon.com/bedrock/latest/userguide/titan-embedding-models.html" }, | ||
| { "text": "Amazon Bedrock prompt caching (native, prefix-based) - complements this pattern", "link": "https://docs.aws.amazon.com/bedrock/latest/userguide/prompt-caching.html" }, | ||
| { "text": "AWS Lambda function URLs", "link": "https://docs.aws.amazon.com/lambda/latest/dg/lambda-urls.html" } | ||
| ] | ||
| }, | ||
| "deploy": { | ||
| "text": [ | ||
| "See the README for the two S3 Vectors setup commands, then: sam build && sam deploy --guided" | ||
| ] | ||
| }, | ||
| "testing": { | ||
| "headline": "Testing", | ||
| "text": [ | ||
| "See the GitHub repo for detailed testing instructions (miss, exact hit, semantic hit, force-invalidate)." | ||
| ] | ||
| }, | ||
| "cleanup": { | ||
| "headline": "Cleanup", | ||
| "text": [ | ||
| "1. Delete the stack: <code>sam delete</code>.", | ||
| "2. Delete the S3 Vectors index and bucket: <code>aws s3vectors delete-index ...</code> then <code>aws s3vectors delete-vector-bucket ...</code>." | ||
| ] | ||
| }, | ||
| "authors": [ | ||
| { | ||
| "name": "Manish S", | ||
| "image": "", | ||
| "bio": "AWS Support Engineer, trying to build things", | ||
| "linkedin": "https://www.linkedin.com/in/manish-s-84199221b", | ||
| "twitter": "" | ||
| } | ||
| ] | ||
| } |
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
The README introduces the full name "AWS Systems Manager" but later uses the abbreviation "SSM Parameter Store". For consistency, prefer "AWS Systems Manager Parameter Store" (or "Systems Manager" after first reference).