Amazon Web Services (AWS) today announced a significant enhancement to Amazon Simple Storage Service (Amazon S3) with the introduction of "annotations," a new metadata capability designed to attach rich, large-scale business context directly to objects. This innovation allows users to store up to 1,000 named annotations per object, with each annotation capable of reaching 1 MB in size, culminating in a total of up to 1 GB of metadata per object. These annotations support flexible formats such as JSON, XML, YAML, or plain text, offering unprecedented versatility. Crucially, annotations can be modified or deleted at any time without requiring the re-writing of the core S3 objects, ensuring that object context remains current and dynamic, a critical feature for modern data management paradigms.
A New Era for Data Understanding and AI-Driven Workflows
This development marks a pivotal moment in how organizations manage and interact with their vast repositories of unstructured data. In an increasingly AI-driven world, the demand for data that is not only stored securely but also deeply understood by autonomous systems is paramount. Organizations are actively building AI agents and sophisticated autonomous workflows that necessitate the ability to find, comprehend, and act on data without constant human intervention. Such agentic workflows demand metadata that can evolve synchronously with the data itself, scale seamlessly across petabytes of objects, and remain readily queryable without incurring expensive retrieval costs. S3 annotations directly address these emerging requirements, bridging the gap between raw data storage and intelligent data utilization.
The traditional challenges associated with managing object metadata have often forced organizations to maintain complex, separate metadata systems or external databases. These external systems frequently lead to synchronization headaches, increased operational overhead, and potential inconsistencies between the data and its descriptive context. By integrating rich metadata directly within S S3, AWS aims to simplify these architectures, enhance data governance, and accelerate the development of advanced data applications.
Technical Capabilities and Flexibility
S3 annotations offer a robust framework for attaching diverse contextual information. For instance, AI-generated transcripts of audio or video files, content ratings for media assets, detailed technical specifications for engineering designs, or compliance attestations for sensitive documents can now reside directly alongside their respective S3 objects. This co-location ensures that the context remains intrinsically linked to the object, automatically moving with it during copy, replication, and cross-region transfers. When an object is deleted, its associated annotations are also automatically removed, maintaining data hygiene and consistency.
A key differentiator is the integration with S3 Metadata. When S3 Metadata is enabled, annotations automatically flow into fully managed annotation tables built on Apache Iceberg. These tables are then readily queryable using powerful analytics engines like Amazon Athena, empowering users to perform complex analytical queries across their metadata at scale. This capability transforms metadata from a passive descriptor into an active, queryable asset, unlocking new possibilities for data discovery and analysis.
Addressing Long-Standing Metadata Challenges
Amazon S3 has long supported various methods for describing objects, each designed for specific purposes. System-defined metadata captures inherent properties like object size, storage class, and creation time, remaining immutable. User-defined metadata allows for small, custom key-value pairs to be added at upload time, but these are also immutable post-upload and limited in size (up to 2 KB). Object tags provide a more flexible option for operational tasks, such as access control, lifecycle management, and cost allocation, supporting up to 10 tags per object, each with limited character length.
While these existing capabilities serve their intended functions effectively, they presented significant limitations when the need arose for attaching much richer, dynamic context without the burden of building and maintaining separate metadata systems. Annotations fundamentally redefine S3’s metadata capabilities by offering a scale and flexibility previously unavailable. Compared to 10 immutable tags or 2 KB of header data, annotations provide mutable, queryable context up to 1 GB per object, distributed across 1,000 individual 1 MB annotations. This vast capacity and mutability are critical for use cases where metadata is not static but evolves over time, such as iterative AI model outputs or ongoing data enrichment processes.

The following table highlights the distinct advantages of annotations:
| Capability | Max size | Mutable? | Best for |
|---|---|---|---|
| System-defined metadata | Fixed | No | Object properties (size, storage class, creation time) |
| User-defined metadata | 2 KB | No | Small custom key-value pairs |
| Object tags | 10 tags, 128/256 chars | Yes | Access control, lifecycle rules, cost allocation |
| Annotations | 1 GB (1,000 × 1 MB) | Yes | Rich business context (JSON, XML, YAML, plain text) |
Historically, the separation of S3 objects from their extensive metadata often led to complex synchronization workflows. These workflows could be costly to develop and maintain, sometimes even exceeding the data storage costs themselves. With S3 Metadata annotation tables, this context becomes inherently queryable at scale through Amazon Athena, eliminating the need for such intricate external systems. Furthermore, AI agents can now discover data through natural language queries using the S3 Tables MCP server, which provides a standardized interface for AI models to interact with and query annotations. This allows for searching objects by their contextual information across any storage class, including S3 Glacier, without the need to restore the objects or incur retrieval charges, a significant cost and time saving.
Getting Started: API and CLI Examples
To begin leveraging S3 annotations, users need to ensure their AWS Identity and Access Management (IAM) policy or bucket policy grants permissions for the s3:PutObjectAnnotation and s3:GetObjectAnnotation actions. Once permissions are in place, annotations can be added to any existing or new S3 object using the PutObjectAnnotation API.
Consider a media company managing a vast library of video assets. They can now attach comprehensive technical specifications and AI-produced summaries directly to each video file.
# Create a JSON file with technical metadata
cat > mediainfo.json << 'EOF'
"codec":"H.265","resolution":"3840x2160","audio_tracks":8,"frame_rate":29.97
EOF
# Attach it as an annotation to a video object
aws s3api put-object-annotation
--bucket my-media-bucket
--key videos/documentary-2026.mp4
--annotation-name mediainfo
--annotation-payload ./mediainfo.json
In parallel, an AI-generated summary can be attached as a separate, plain-text annotation:
# Attach a plain-text AI-generated summary as a separate annotation
echo "A 90-minute nature documentary covering wildlife migration patterns across three continents, featuring aerial footage and underwater sequences. Languages: English, Spanish, Portuguese." > ai_summary.txt
aws s3api put-object-annotation
--bucket my-media-bucket
--key videos/documentary-2026.mp4
--annotation-name ai_summary
--annotation-payload ./ai_summary.txt
These commands illustrate how two distinct annotations—mediainfo (structured JSON) and ai_summary (plain text)—can be associated with the same video object. Each annotation is uniquely named, allowing for independent reading, modification, and deletion. This naming convention facilitates concurrent enrichment workflows; for example, one team can add technical metadata while another adds content classifications without conflict.
Retrieving a specific annotation is straightforward using the GetObjectAnnotation API:
aws s3api get-object-annotation
--bucket my-media-bucket
--key videos/documentary-2026.mp4
--annotation-name mediainfo
./mediainfo-output.json
To list all annotations associated with an object, the ListObjectAnnotations API provides a comprehensive overview:
aws s3api list-object-annotations
--bucket my-media-bucket
--key videos/documentary-2026.mp4
When an annotation is no longer required, it can be precisely removed using the DeleteObjectAnnotation API:
aws s3api delete-object-annotation
--bucket my-media-bucket
--key videos/documentary-2026.mp4
--annotation-name mediainfo
Existing annotations can be updated at any time by simply calling PutObjectAnnotation again with the same annotation name. For large objects uploaded using multipart upload, annotations can be attached after the multipart upload is complete, ensuring flexibility in the data ingestion process.

Querying Annotations at Scale with S3 Metadata Tables
While attaching annotations to individual objects is powerful, the true transformative potential lies in the ability to query across all annotations at scale. This is achieved by enabling S3 Metadata annotation tables on an S3 bucket. Upon activation, S3 automatically indexes all annotations into a fully managed Apache Iceberg table, referred to as an annotation table. These tables are designed for high-performance querying and are compatible with Amazon Athena and any other Iceberg-compatible engine.
Enabling annotation tables is done via the S3 console or the CreateBucketMetadataConfiguration API. An example configuration might look like this:
"JournalTableConfiguration":
"RecordExpiration": "Expiration": "DISABLED"
,
"InventoryTableConfiguration": "ConfigurationState": "DISABLED" ,
"AnnotationTableConfiguration":
"ConfigurationState": "ENABLED",
"Role": "arn:aws:iam::123456789012:role/S3MetadataAnnotationRole"
This configuration instructs S3 to automatically capture all annotations within the specified bucket into a queryable table. Once applied, any new annotation attached to objects in this bucket will appear in the table within approximately one hour. For buckets that already have an existing metadata configuration, the UpdateBucketMetadataAnnotationTableConfiguration API can be used to enable the annotation table feature.
A key advantage of S3 annotation tables is their schema flexibility. Unlike traditional metadata tables that often require predefined schemas, annotation tables automatically adapt to any JSON, XML, or YAML structure provided. Each annotation becomes a row in the table, with its content stored in a text_value column, enabling seamless querying across diverse annotation structures without the need for complex schema migrations.
For buckets with pre-existing annotated objects, S3 automatically initiates a backfill process to populate the annotation table. This background process ensures that all historical metadata is captured, though it may take several hours to days depending on the volume of objects.
Advanced Querying and AI Integration
The integration with Amazon Athena allows for sophisticated SQL queries against the annotation tables. For instance, to identify all video assets with more than eight audio tracks across an entire bucket, a query like this can be executed:
SELECT DISTINCT bucket, object_key
FROM "s3tablescatalog/aws-s3"."b_my_media_bucket"."annotation"
WHERE name = 'mediainfo'
AND CAST(json_extract_scalar(text_value, '$.audio_tracks') AS INTEGER) > 8
This query efficiently scans the annotation table for annotations named mediainfo, extracts the audio_tracks field from the JSON content, and returns the unique object keys where the audio track count exceeds eight.
For near real-time tracking of annotation changes, the journal table (another component of S3 Metadata) can be utilized:
SELECT bucket, key, version_id, record_timestamp, annotation.name
FROM "s3tablescatalog/aws-s3"."b_my_media_bucket"."journal"
WHERE record_timestamp >= (current_date - interval '1' day)
AND annotation.name IS NOT NULL
AND record_type IN ('CREATE_ANNOTATION', 'DELETE_ANNOTATION')
This query identifies all objects that received new annotations or had annotations deleted within the last 24 hours, making it ideal for building event-driven workflows that react dynamically to metadata changes.

The potential for AI integration is particularly significant. By leveraging agents in Amazon SageMaker Unified Studio or any IDE equipped with the S3 Tables MCP server, users can employ natural language to search objects based on their annotations. Imagine asking, "find all PG-rated movies with Spanish subtitles from 2023." Such a query, which might have previously required hours of querying disparate systems, can now yield results in seconds, dramatically enhancing data discoverability for AI models and human analysts alike. This capability moves AWS closer to a vision of truly intelligent data lakes where data is not just stored but actively understood and utilized by AI.
Broader Implications and Market Impact
The introduction of S3 annotations underscores AWS’s commitment to evolving S3 beyond a mere storage service into a foundational platform for intelligent data management. This feature positions S3 more strongly in a competitive cloud storage market by directly addressing the increasing complexity of data workflows, especially those driven by artificial intelligence and machine learning. As data volumes continue to explode and the sophistication of analytical and generative AI models grows, the ability to embed rich, queryable context directly within storage objects becomes a critical differentiator.
This advancement is expected to accelerate innovation across various industries. In media and entertainment, it simplifies content cataloging, rights management, and automated content processing. In healthcare, it can facilitate better management of patient records, imaging data, and research datasets with embedded compliance and contextual information. Financial services can leverage annotations for regulatory compliance, audit trails, and risk assessment by attaching granular context to transactional data. Scientific research and genomics will benefit from the ability to link experimental parameters, data provenance, and analytical results directly to raw datasets.
Availability and Pricing
Amazon S3 annotations are available today across all AWS Regions, including the AWS China Regions, providing global accessibility for this powerful new feature. Annotation tables are available in all AWS Regions where S3 Metadata is already supported.
Regarding pricing, annotation storage is consistently billed at S3 Standard rates, regardless of the parent object’s storage class (e.g., S3 Glacier). This simplified pricing model ensures transparency and predictability for users. Full pricing details are available on the Amazon S3 pricing page.
This strategic release from AWS empowers organizations to manage petabytes of data with unprecedented contextual depth and agility. Whether the goal is to build autonomous AI agents that discover data independently, manage complex media assets with evolving metadata, or track compliance contexts for archived datasets, S3 annotations provide the necessary scale, flexibility, and integration to attach rich metadata directly to objects, eliminating the need for cumbersome, separate systems. This represents a significant step forward in the journey towards fully intelligent and self-managing data ecosystems in the cloud.
