Conversation
* fix: byte-order handling for structured dtypes in the bytes codec The bytes codec neither byte-swapped structured-dtype fields to its configured endian on encode (numpy reports byteorder '|' for void dtypes, so the top-level byteorder comparison never detected a mismatch) nor honored its endian when decoding, silently corrupting any structured data whose field byte order differed from the stored one (e.g. virtual references to external big-endian data). Encode now detects byte-order mismatches by comparing full dtypes via newbyteorder, and decode reinterprets raw bytes in the stored byte order before converting to the data type's declared byte order, so the stored layout (codec state) and the in-memory layout (array data type) are independent. Closes zarr-developers#4141 Assisted-by: ClaudeCode:claude-fable-5 * test: fold structured byte-order cases into existing bytes codec tests Extend test_endian's parametrization with structured dtypes and test_bytes_codec_sync_roundtrip with endian/dtype parametrization plus stored-layout and decoded-dtype assertions, instead of adding parallel test functions for the same properties. Assisted-by: ClaudeCode:claude-fable-5 * refactor: rename stored_dtype to view_dtype in BytesCodec decode The variable is the dtype used to view the raw chunk bytes (byte order from the codec's endian configuration), not a property of the stored data or of the returned buffer, which always carries the array's declared dtype. Assisted-by: ClaudeCode:claude-fable-5 * docs: note that the decode-side byte-order conversion copies the chunk Assisted-by: ClaudeCode:claude-fable-5
…etails Add pre-release clarifications to the 3.3.0 section: - zarr-developers#4141: note the read-time behavior change for structured-dtype arrays written by <= 3.2.1 with an explicitly non-native bytes-codec endian. - zarr-developers#3417: note that V3 metadata omitting endian for multi-byte dtypes now fails at open time with ValueError instead of assuming native order. - New Misc entry: typing_extensions >= 4.14 is now required (Sentinel), and the gpu extra skips CuPy on macOS. - zarr-developers#3963 / zarr-developers#3968: clarify that enum-specific idioms on the now-string attributes raise immediately (no deprecation cycle), while equality against enum members still holds. Assisted-by: ClaudeCode:claude-fable-5
Assisted-by: Codex:GPT-6
Assisted-by: Codex:GPT-6
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
🤖 AI text below 🤖
Amends the 3.3.0 release notes for structured byte-order handling, missing bytes-codec endian metadata, dependency requirements, and deprecated enum usage.
The compatibility note now describes the actual 3.2.1 code: structured decoding ignored the codec's byte order, and the encoder treated aggregate structured byte order as native when deciding whether to convert. Upgrading can change legacy reads depending on stored bytes and field dtypes. A default/native codec setting alone does not prove a mixed-endian array is unaffected; the previous blanket assurance and claim that all non-native configurations stored unswapped data were incorrect.
Missing endian is also not universally rejected: the structured legacy path warns and assumes little endian, while other multi-byte dtypes require the configuration. It is not established that only third-party writers can produce affected metadata.
The other notes clarify the typing_extensions floor, macOS GPU-extra dependency condition, and enum-specific idioms that stop working. Documentation-only factual follow-up; applicable commit hooks pass. Earlier empirical enum checks are historical validation, not a new full release audit.