What Is S3’s Default Multipart Checksum Design?
Amazon S3 automatically protects newly created objects with a data integrity checksum, including multipart uploads, using CRC64NVME by default. AWS announced the change for new objects in January 2024, following work in 2023 that added support for additional checksum algorithms and checksum modes. CRC64NVME produces a 64-bit value based on the object contents, so S3 can detect corruption caused by the network, storage hardware, software defects, or incomplete processing.
Also worth reading: How Should a Cross-Cloud Object-Storage Checksum Policy Work in 2026? · How Can Teams Verify Objects After Migrating Data Across Cloud Object Stores? · How Should Platform Teams Design a Multi-Cloud Storage Architecture in 2026?
The design is not a replacement for TLS, object versioning, replication, or application-level validation. It is a compact data-plane control that lets S3 and its client compare evidence that the received bytes match the bytes that were sent. For a multipart object, S3 must be able to connect the checksums calculated by individual parts without calculating an ordinary full-object checksum on every upload path.
The important terminology is “new objects.” Enabling the default does not retroactively calculate checksums for every existing object or trigger a background rewrite. Objects already stored before the change generally retain their previous checksum behavior unless copied or otherwise rewritten. Applications should therefore treat the default as forward protection, not as a fleet-wide integrity retrofit.
CRC64NVME offers a stronger raw detection range than the older CRC32 family, which uses 32 bits. Under a simple uniform-error model, CRC64 has roughly a 1-in-2^64 probability of accepting one corrupted byte sequence as matching, compared with about 1-in-2^32 for CRC32. A checksum is not encryption, and a deliberately engineered collision remains possible; the practical value comes from making accidental corruption much less likely to pass silently.
How Checksum Validation Works for Multipart Uploads
A multipart upload splits a large object into parts, uploads those parts independently, and then asks S3 to assemble them into one object. The classic limits include a maximum of 10,000 parts per multipart upload, a minimum part size of 5 MiB except for the final part, and a maximum object size of 5 TiB for S3 Standard and S3 Standard-IA. These limits influence checksum behavior because S3 can combine the checksum result of each part into a final result without retransmitting the object.
For checksums designed to support this workflow, AWS uses a composite form. The upload client calculates a checksum for each part, and S3 records that result with the corresponding uploaded part. S3 derives an object-level value from the ordered collection of part checksums, upload IDs, and part numbers. The order matters because a valid object cannot simply reorder its parts while retaining the same multipart checksum result.
S3 can return the final checksum through response headers and API operations, while the client can supply its own checksum or an explicitly selected algorithm and algorithm family. The service then compares the transmitted, S3-calculated, and, where applicable, final values. This approach avoids a second full read merely to recalculate the object checksum after completion, which would add time, request capacity, and potentially data-transfer charges.
That efficiency is especially useful for datasets measured in terabytes, although the client must still calculate part checksums while streaming the data. If the source is already local, the operation is primarily CPU and memory work. If a transfer service is uploading directly from another cloud, checksum calculation can be combined with streaming, but doing it correctly may require temporary buffering when an algorithm or part boundary cannot otherwise be met.
Why AWS Selected CRC64NVME and Added Other Algorithms
The default was not intended to make every checksum algorithm obsolete. AWS added support for CRC32C, CRC32, SHA-1, and SHA-256, while making CRC64NVME the default for newly created objects. CRC32C is common in data-center and storage systems, CRC32 remains inexpensive and widely supported, and SHA-family algorithms provide familiar options for security-oriented or compliance-oriented workflows.
CRC64NVME is designed for high-throughput storage workloads and offers a 64-bit result without the cost of a cryptographic hash in the common default path. SHA-256 produces a 256-bit digest, which can reduce the practical probability of an accidental or malicious collision, but it normally requires more CPU work and is not automatically an authenticity mechanism. A client must not confuse a stored SHA-256 checksum with a digital signature unless the signature is separately created and verified.
AWS also distinguishes full-object checksums from composite checksums. A full-object checksum represents all bytes in the completed object. A composite checksum represents the ordered part-level results used during multipart completion. Consequently, a value calculated by a different checksum mode or algorithm is not directly comparable, even when both values are labeled SHA-256 or CRC.
| Feature | Default multipart behavior | Alternative checksum configuration |
|---|---|---|
| Default algorithm | CRC64NVME for new objects | CRC32, CRC32C, SHA-1, or SHA-256 can be selected where supported |
| Digest width | 64 bits | 32 bits for CRC32 and CRC32C; 160 bits for SHA-1; 256 bits for SHA-256 |
| Multipart representation | Composite checksum derived from ordered part results | Same general composite principle, with different digest values and calculation costs |
| Automatic validation | Applies to newly uploaded S3 objects | Requires the client or bucket/workflow to request the algorithm and family correctly |
| Malicious tamper protection | Not a cryptographic guarantee | SHA-family digests are stronger integrity primitives but still require protected metadata or signatures for authenticity |
| Legacy objects | Not retroactively rewritten | Must be copied, uploaded again, or explicitly processed if validation is required |
| Additional charge | No separate checksum price in ordinary S3 usage | No separate checksum price, but the selected operation and storage class can incur normal charges |
For a new S3 integration, the safest baseline is to accept the CRC64NVME default and verify that the SDK or transfer utility preserves the checksum metadata returned by S3. Do not disable checksum headers or validation merely to simplify a migration. Most managed SDKs negotiate supported checksum settings, but custom HTTP clients, gateways, and multipart libraries can behave differently.
A team operating a cross-cloud object-storage SaaS should test the complete path rather than only a direct SDK call. That path may include a customer’s virtual private network, an AWS Direct Connect or internet connection, a proxy, a data-ingestion worker, an S3-compatible endpoint, and a downstream analytics system. The test should confirm that the part size, checksum algorithm, checksum family, part order, and final completion request remain consistent.
Bucket policies can restrict actions and encryption settings, but checksum support is primarily an API and client concern. Teams should also capture the final object checksum in a controlled metadata or audit system if downstream verification is required. Copying that value into an unprotected, user-writable location can make the verification record no more trustworthy than the object it is intended to protect.
The practical validation test should deliberately introduce an error. Upload an object, alter one local part before sending it without updating its expected checksum, and confirm that S3 or the transfer tool rejects the mismatch. Then repeat the test with a valid upload, including one multipart object with at least two parts. This negative test matters because a workflow that reports a checksum successfully but does not compare it has not implemented meaningful validation.
Before a production rollout, teams should identify applications that replay multipart completion XML or copy raw upload records. Stored part metadata is part of the multipart protocol, and an obsolete client may assume a single full-object checksum where S3 now returns a composite value. Compatibility testing is more reliable than assuming that a major AWS API change will break every client, but it is still necessary for custom implementations.
Cost, Performance, and Storage-Class Considerations
S3 checksums do not carry a separate line-item charge. Normal request, data-transfer, retrieval, and storage charges still apply. Uploading a 10 TiB object with checksum calculation does not become free because the service validates its bytes, and duplicating an object solely to add a checksum can create meaningful PUT, storage, and lifecycle costs.
The computational cost differs by algorithm. CRC64NVME was selected partly because it balances detection strength with throughput. CRC32 and CRC32C are also efficient, while SHA-256 usually demands more CPU per byte. In many network-bound transfers, the performance difference is small; in local uploads, high-speed capture, or compute-constrained migration workers, it can become visible.
Storage class does not change the purpose of the checksum, although it changes related economics. S3 Standard-IA and S3 Glacier Instant Retrieval have minimum object sizes and retrieval-duration considerations, while S3 Glacier Flexible Retrieval and S3 Glacier Deep Archive require asynchronous restoration before most data can be read. Checksums help validate restored or downloaded bytes, but they do not make retrieval immediate.
For repeated migration or disaster-recovery testing, a provider or customer may choose to store copies in more than one cloud. That introduces a separate data-consistency requirement: the checksum identifies content, but it does not prove that the same object version, metadata, legal hold, or retention policy exists in each cloud. Cross-cloud replication systems should record a content identifier, the algorithm, the checksum family, and the source version alongside the destination receipt.
A 5 TiB maximum object is large enough that copying or re-uploading solely to obtain a checksum can be expensive in time and egress. Whether rewriting is justified depends on the value and risk of the dataset, the object’s age, the existing protection model, and the cost of any corruption not yet detected. Low-value, regenerable test data may not justify that work, while regulated, archival, or hard-to-recreate records may justify explicit verification.
Common Mistakes and Compatibility Traps
The first common mistake is treating a checksum as proof that the source was correct before upload. S3 can confirm that the received bytes correspond to the expected checksum, but it cannot know that the source file was already wrong. A second mistake is comparing values without recording whether the checksum was full-object or composite. Two SHA-256 values can be different because they use different multipart checksum families, not because either service failed.
Another trap is changing the part order or reusing a completed part record. Multipart composition is ordered. If a tool copies parts between uploads, changes their sequence, or mixes uploads created with incompatible settings, the resulting checksum can be rejected or misinterpreted. The completion request also needs every intended part and its correct ascending part number; a successful individual part upload does not by itself mean the assembled object was successfully created.
Some operators disable automatic checksum behavior because a gateway reports an unsupported header. That response may indicate a custom client, an older S3-compatible implementation, or a missing dependency rather than a need to abandon validation. Updating the client or transfer utility is usually preferable, but teams must not assume that another cloud’s “S3-compatible” endpoint implements AWS checksum semantics identically.
A final mistake is assuming that SHA-256 automatically protects against a malicious actor who can replace both an object and its stored checksum. Ordinary checksums detect accidental damage and many transmission faults. Stronger protection against intentional substitution requires an independent trust anchor, such as a digitally signed manifest, a key-management-controlled checksum record, or an immutable audit log.
When to Act, Re-evaluate, or Avoid an Immediate Rewrite
New workloads should act now by using current S3 SDKs, testing multipart uploads, and preserving the checksum response in their operational records. There is no reason to bypass the default merely to reproduce legacy behavior on a new bucket. Teams can enable a deliberate algorithm choice when a regulatory framework, existing data pipeline, or interoperability requirement calls for SHA-256 or another supported option.
Existing workloads should first inventory their upload mechanisms. Managed SDKs, AWS CLI, Storage Gateway, Snowball, and third-party migration tools may have different checksum behavior, and no single conclusion applies to all of them. The inventory should identify which systems create new writes, which call CompleteMultipartUpload, and which independently recompute checksums after transfer.
A blanket rewrite is rarely the first action. It doubles or otherwise increases object-processing activity, can trigger lifecycle and replication consequences, and may not fix the real issue if corruption is occurring above the S3 API. Instead, prioritize systems with large business impact, weak retry logic, long retention periods, or expensive reproducibility. Objects with short retention and available pristine sources may have a lower expected benefit from immediate recomputation.
Organizations should re-evaluate their policy when SDKs change checksum defaults, when a cloud provider modifies compatibility behavior, or when a new integrity requirement distinguishes accidental corruption from malicious tampering. They should also revisit part sizing. Larger parts can reduce request and completion overhead, but an excessively large part increases retry cost after failure; smaller parts improve retry granularity but increase request counts and multipart bookkeeping.
The balanced 2026 recommendation is therefore straightforward: use S3’s default CRC64NVME protection for new objects, select a stronger digest only when there is a concrete reason, and verify that the entire application path preserves and compares the correct checksum. Do not confuse new default coverage with retroactive validation, and do not pay to rewrite the entire bucket without a risk-based reason.
How Cross-Cloud OSS Platforms Can Use the Design Responsibly
For a B2B cross-cloud data plane, S3 multipart checksums should be treated as one layer in a portable integrity contract. The contract can name the source provider, bucket or tenant boundary, object version, byte size, checksum algorithm, checksum family, and final checksum. A destination adapter can then verify that its received object matches the same declared digest while recognizing that a multipart composite checksum is an S3 protocol value rather than a universal cloud-wide object identifier.
Content-addressable systems may calculate a full SHA-256 independently at the destination. That can be useful for deduplication and cross-cloud portability, but it is not numerically identical to S3’s multipart SHA-256 composite value. The system should store both when necessary: the provider-native checksum for protocol validation and a separate content digest for neutral manifest or deduplication purposes.
Platform teams should expose checksum status without implying perfect trust. A customer interface might say “integrity verified against the supplied digest,” not “the object is authentic.” If the digest came from the same compromised upload client as the data, verification proves internal consistency rather than identity. For stronger assurance, sign a manifest containing object identifiers, versions, sizes, and digests using a key held outside the upload worker.
Migration tooling should also make failures actionable. A checksum mismatch needs the failed part, request ID, retry count, source and destination endpoints, and a clear distinction between a transient transfer problem and a protocol mismatch. Automatic retries can resolve network corruption, but repeatedly retrying an incompatible algorithm or family creates delay without improving correctness. Alert thresholds should therefore depend on failure rate, data volume, and customer impact rather than treating every mismatch as an isolated network event.
This approach allows an OSS platform to benefit from S3’s managed default without falsely advertising identical behavior across every provider. It also preserves room to map AWS CRC64NVME, another provider’s native checksum, and a platform-owned SHA-256 manifest into separate fields. That separation is more reliable than flattening several integrity models into one ambiguous “verified” flag.