Conversation
Slice the byte vector before converting to str, avoiding a whole-buffer conversion while compaction may leave invalid UTF-8.
Collaborator
|
Thanks for the pull request, and welcome! The Rust Project has assigned @JohnTitor (or someone else) to review your changes, you should hear from them (or someone else) within the next two weeks. Please see the contribution instructions and our LLM policy for more information. Why was this reviewer chosen?The reviewer was selected based on:
|
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
During compaction,
String::retainreads the next character withg.s.get_unchecked(read..len). This invokesString’sDeref, constructing a&strover the entire0..lenbuffer before selecting the unread tail.For example, removing
'é'from"éab"and moving'a'left changes the buffer fromC3 A9 61 62to61 A9 61 62. The next iteration reads the valid'b'tail, but first converts the entire buffer, including the orphanedA9continuation byte, to&str.Constructing a non-UTF-8
&stris not immediate language-level UB under the Reference’s validity requirements. However, this conversion violatesfrom_utf8_unchecked’s documented safety precondition. The libs discussion in #134598 is defining which operations accept invalid UTF-8;str::get_uncheckedis absent from the tentative list.The fix slices
read..lendirectly from the byte vector and converts only that untouched tail to&str. The original string was valid UTF-8, andreadremains on a character boundary. #154583 uses the same approach forextract_if.Behavior is unchanged. A separate codegen comparison produced identical optimized x86-64 assembly for the tested predicates.
Validation
I ran:
./x test library/alloc: all 1,491 tests passed, includingtest_retain../x test tidy: passed../x miri library/alloc:test_retainand theretaindocumentation tests passed under both Stacked Borrows and Tree Borrows, without reported UB.I also ran differential comparisons covering:
"éab"case.These total 2,094,099 comparison cases; the panic cases are included in the exhaustive count.
AI involvement
An AI assistant initially noticed the potential issue during a verification attempt. I independently verified the finding and confirmed the issue myself. I wrote the patch, SAFETY comment, and this description.
I used AI assistance to research the implementation and related PRs, and to assist with differential comparisons and Miri checks.
Related: #150067, #134598, #154583.