MIR move elimination - #3943
MIR move elimination#3943Amanieu wants to merge 16 commits into
Conversation
|
For It seems to be extremely strange to say I need a |
We can drop The purpose of a |
Wouldn't it make more sense to have something similar to |
|
Calling let x = 1;
let y = &x;
drop(move x);
let z = *y; // Fails because x was moved. Removing `move` fixes this.Anyways, |
|
|
|
||
| The proposed behavior of freeing a local variable's allocation on move only applies when the entire variable is moved. This is not the case when only a part of the variable is moved (e.g. only one field of a struct) because a re-initialized field must retain the address it had before, reintroducing the same NB issue. | ||
|
|
||
| Even in the case where all of the fields of a local variable have been moved out one-by-one, the local will not be freed. |
There was a problem hiding this comment.
Even in the case where all of the fields of a local variable have been moved out one-by-one, the local will not be freed.
Why not?
|
|
||
| Even in the case where all of the fields of a local variable have been moved out one-by-one, the local will not be freed. | ||
|
|
||
| With that said, we would like to keep the door open for potentially switching to operational semantics with NB in the future. So although the proposed opsem does not consider accessing a moved field as UB, we would like users to avoid relying on this behavior since it may change in the future. |
There was a problem hiding this comment.
Why can't we say that accessing the moved field is UB? Because we don't have NB, the compiler can't exploit that UB by stashing another allocation with an observable address in the empty space. But that doesn't mean it can't still be UB detected by Miri, if we want users to avoid relying on it. And the compiler could even make use of the UB, to stash an allocation whose address it can prove is never observed.
There was a problem hiding this comment.
There's no reason we can't do it, it's just that I am not doing so in this RFC and instead leaving to future work (I will update the future possibilities section with this). There are 2 main reasons for this:
-
Adding support for partial moves makes the opsem (and by extension Miri) much more complex since we now needs to track which bytes of a local have been moved out and become "inactive" (UB to access). What happens to the padding of a struct if one field is moved out? What happens to the discriminant of an enum like
Option<T>if theSomevalue has been moved out and the layout is optimized (the discriminant occupied the same bytes as the value)? These are all questions that would have to be answered. -
The proposed MIR optimization pass can't easily take advantage of this, and even if it could (while respecting address observation rules) then I expect the benefit over the existing proposed pass would be minimal. It's just not worth the extra complexity of tracking lifetimes separately for every field of a local.
|
|
||
| ## Drawbacks | ||
|
|
||
| ### `Copy` is no longer "free" |
There was a problem hiding this comment.
This makes it sound as though types with Copy would be less efficient than with the status quo, but as far as I understand this is not the case. Copy types would at worst be as efficient as with the status quo.
This sounds more like a limitation (AKA an opportunity for future work) than a drawback.
It's also not clear to me why the mere act of implementing Copy for a type would inhibit the optimization in practice. Surely code that was originally written for a non-Copy type would have an access pattern where every copy could be optimized into a move?
There was a problem hiding this comment.
This makes it sound as though types with
Copywould be less efficient than with the status quo, but as far as I understand this is not the case.Copytypes would at worst be as efficient as with the status quo.
That is correct.
This sounds more like a limitation (AKA an opportunity for future work) than a drawback.
Sure, I can move this to the future work section.
It's also not clear to me why the mere act of implementing
Copyfor a type would inhibit the optimization in practice. Surely code that was originally written for a non-Copytype would have a copy pattern where every copy could be optimized into a move?
It inhibits the optimization if the value has been borrowed. A move invalidates borrows whereas a copy doesn't. I've update the text with an example.
There was a problem hiding this comment.
Finally had the time for a first read over this.
I haven't yet had the chance to look at your MiniRust patch, that may answer some of the questions that came up here. But of course ideally the RFC itself is already crystal clear about the intended MR semantics. :)
|
|
||
| #### Initialization | ||
|
|
||
| `StorageLive` no longer allocates the underlying memory for a local. Instead, any MIR statement or terminator which writes to a place that has no `Deref` projections[^2] will implicitly allocate the storage for that local[^3] before writing to it. This has no effect if the storage for that local is already allocated[^4]. |
There was a problem hiding this comment.
To make it sound to generate LLVM lifetime markers from StorageLive, we will need to do something in the opsem for StorageLive.
Reading on, you clarify this later, but the order in which you explain this is confusing.
There was a problem hiding this comment.
I've kept the original ordering since for most people it's better to give an overview of the intended semantics first before diving into a more formal specification for StorageDead/StorageLive. However I have added a link to that in the section intro.
|
|
||
| `move` operands only have the effect of de-allocating the storage of a local when used with a bare, unprojected local. If the local has projections then `move` behaves identically to `copy`. | ||
|
|
||
| #### New semantics of `StorageLive` and `StorageDead` |
There was a problem hiding this comment.
This RFC also inevitable deeply changes the semantics of assignments. There should be a section on that.
And this section is where I have the biggest disagreement with the RFC: the RFC proposes to make (de)allocation of locals an implicit side-effect of (some) place/value expressions. I think that's a problem for multiple reasons
- It breaks the property that expression evaluation can be arbitrarily reordered with each other, making it harder to reason about MIR.
- For places it's not even clear how evaluation should work. I can make guesses but the RFC is not very clear about it. (EDIT: This one is answered by the MiniRust patch.)
I would propose that instead the (de)allocation happens as part of the assignment operator. The way I think about it is that the operator performs the following steps:
- evaluate the value expression (RHS)
- deallocate a set of locals
- allocate a set of locals
- evaluate the place expression (LHS)
- store the value into the place
In MiniRust, it's probably easiest to just annotate those two sets of locals as part of the syntax of assignment itself. In MIR, we might want to make that implicit. That would then be able to capture the syntactic ruls you have proposed elsewhere:
- for the locals to deallocate: if the RHS is
Use(Move(local)), thenlocalis deallocated, otherwise nothing gets deallocated. (Or is it really allMove(local)operands on the RHS? I can see it being useful for aggregate initialization but I doubt this makes much sense for binops.) - for the locals to allocate: if the LHS is
local.non-deref-projs, thenlocalgets allocated, otherwise nothing gets allocated
There was a problem hiding this comment.
As per our discussion on Zulip, I've kept the existing semantics of doing allocation/de-allocation as part of destination place/input operand evaluation. I've added bulk allocation/deallocation as an alternative option, but in the end both options are workable for MIR optimizations, the only difference is in how they are handled in Miri/MiniRust.
It is outside the scope of this RFC and just confuses things.
This comment was marked as outdated.
This comment was marked as outdated.
|
|
||
| This RFC proposes to instead only have the allocations of variables be live while they are *initialized*. This means that the underlying memory for variables is: | ||
| - allocated at the point where it is initialized (instead of where it is declared). | ||
| - freed when a variable of a non-`Copy` type is *moved* (or at the end of its scope, whichever is first). |
There was a problem hiding this comment.
How does this work with unwinding (which causes multiple variables to be deallocated "at the same time") from a scope with variables that don't have destructors? See rust-lang/rust#147875
There was a problem hiding this comment.
It just works? The lifetime of a local on the unwind path ends either at StorageDead or at an UnwindResume/UnwindTerminate terminator.
|
We happened to look at this in an @rust-lang/lang RFC review. We're currently assuming that the RFC qua RFC doesn't need attention from lang until the experiment has progressed further, and that you'll nominate it when you need review here. |
|
Right, the current plan is to first land rust-lang/rust#157943 under the |
There was a problem hiding this comment.
Looks very good overall. :)
I only carefully reviewed the part relevant for MIR semantics. There's a big section on how to do the MIR optimizations that I skipped, but also I don't think the details of that section should be considered normative anyway.
Co-authored-by: Ralf Jung <post@ralfj.de>
|
I implemented Miri support for the new semantics in rust-lang/rust#162048. Two details in particular stood out. ZST handlingThere are several places in MIR intrinsic lowering that assume that unit types (and in one case ZST closures) don't need to be initialized before being accessed, which is not allowed by the new semantics. This is because those ZSTs don't have an address yet (which requires an allocation). Additionally the The current solution is to keep the old semantics for ZSTs by having them be fully allocated by Temporary allocationsEvaluating operands before the destination place in assignments is implemented by creating a temporary allocation with the same type as the destination place, evaluating the full rvalue into that and then evaluating the destination place and copying the result from the temporary allocation into that. This temporary allocation shouldn't be observable in practice since there is no pointer with a valid tag to access it. However it may cause Miri to be slower (I haven't benchmarked it). Temporary allocations are also used for moved-out locals and moved call arguments. |
|
I've updated the RFC based on the experience from implementing Miri support for these semantics and then validating the existing MIR passes against it. Main changes are:
|
…-to-move, r=tmiasko MIR move elimination [2/6]: TailCopyToMove Depends on rust-lang#163335 This PR implements a pre-pass for the `MoveElimination` pass from rust-lang/rfcs#3943. `TailCopyToMove` turns `Copy` into `Move` just before a `Return` terminator. This is valid even for places that have been borrowed because the `Return` will invalidate borrows anyway. This is necessary to allow the source of the copy to be unified with the return place. r? tmiasko
|
|
||
| [^dest-alias]: The normal rules would allow this overlap, which assignments use. But this would break in-place argument passing in codegen. | ||
|
|
||
| These rules continue to allow codegen to pass `move` arguments in-place: the deallocation of bare-local `move` operands allows the corresponding argument in the callee to re-use the same address as the deallocated local. Also note that the exceptions are specific to `Call` terminators. They do not apply to `TailCall`, `InlineAsm`, `Yield`, which continue to use the standard operand evaluation rules. |
There was a problem hiding this comment.
They do not apply to
TailCall,InlineAsm,Yield, which continue to use the standard operand evaluation rules.
Those rules are for statements. I have no idea what it means to apply them to terminators.
Call and TailCall having different rules sound quite confusing and is IMO undesirable.
There was a problem hiding this comment.
There's nothing specific about these rules that make them only apply to statements?
Call and TailCall already have different rules today because TailCall doesn't allow in-place passing. In fact of all terminators, Call is the special case because it is the only MIR statement/termiantor that treats move with special semantics to allow in-place calls. I've previously argued that we should use something other than Operand for call terminators for this reason.
| This RFC makes calls mostly follow the same rules as assignments: operands are evaluated left to right, with bare-local moves deallocating the local as part of evaluating that operand. The destination place is evaluated last and, if direct, will allocate its base local if needed. | ||
|
|
||
| There are 2 exceptions to this: | ||
| - `move` operands with projections keep their old semantics: the moved place is donated to the callee. This place is required to not overlap with the destination place. |
There was a problem hiding this comment.
So move _1 and move _1.field have completely different semantics? That doesn't sound like a good idea, it sounds like a headache and a footgun.
There was a problem hiding this comment.
That's already the case for normal (non-call) operand evaluation: move _1 will de-allocate the local but move _1.field is treated the same as a copy.
There was a problem hiding this comment.
move *_1 is definitively treated as move, not copy (otherwise impl FnOnce() for Box<dyn FnOnce()> couldn't work). And cg_clif also treats move _1.field as move and doesn't copy it on the caller side, allowing the callee to reuse it in-place: https://github.com/rust-lang/rust/blob/7d2cd0fbc092625ea371da704f15f80216d2220e/compiler/rustc_codegen_cranelift/src/abi/mod.rs#L416-L424
There was a problem hiding this comment.
That's already the case for normal (non-call) operand evaluation
Before this RFC it's not the case, it's all nice and homogeneous.
If this RFC makes things messy also for non-call operands that's not really helping.^^
There was a problem hiding this comment.
The main change in this RFC is that moves of a whole local now deallocate the local. Otherwise it behaves the same as before.
| Old semantics | New semantics | |
|---|---|---|
Normal operand copy _1 |
Copies the value, source unchanged | Copies the value, source unchanged |
Normal operand move _1 |
Same as copy | Deallocates _1 during evaluation |
Normal operand move _1.field |
Same as copy | Same as copy |
Call operand copy _1 |
Copies the value, source unchanged | Copies the value, source unchanged |
Call operand move _1 |
Protects _1 for duration of call |
Deallocates _1 during evaluation |
Call operand move _1.field |
Protects _1.field for duration of call |
Protects _1.field for duration of call |
(.field in this table generalizes to any projection, including deref)
Note that both deallocation and protection permit the codegen optimization of in-place passing. Deallocation allows it because it allows the local in the callee's stack frame to re-use the newly freed address. Protection allows it because the protected area is temporarily donated for the callee to allocate its locals from. No changes to codegen were required by the change from protection to deallocation.
There was a problem hiding this comment.
Deallocation also cleanly detects overlaps with other move operands and the destination as UB:
- UB if the local was already protected or deallocated by an earlier operand evaluation.
- Later attempts to protect or deallocate the place will be UB since it has been deallocated.
And finally, it deallocates moved locals before the callee executes, which means the stack frame doesn't need to keep track of a set of locals to deallocate on return.
| - We specifically forbid the destination place from having the same base local as a bare-local moved argument.[^dest-alias] | ||
|
|
||
| [^dest-alias]: The normal rules would allow this overlap, which assignments use. But this would break in-place argument passing in codegen. | ||
|
|
There was a problem hiding this comment.
What does "forbid" mean? Is it UB, or invalid MIR?
This all feels like it's papering over a deeper problem by adding increasingly baroque exceptions and special cases. :/
There was a problem hiding this comment.
Invalid MIR. It was already UB in the old semantics because of the overlap between source and destination, this just makes the guarantee stronger.
There was a problem hiding this comment.
Previously UB naturally fell out of the semantics without special cases. Now there's more special cases that everyone dealing with MIR semantics needs to figure out and remember. That's an increase in entropy.
Why is it invalid MIR now?
There was a problem hiding this comment.
Because without that extra rule, it would be legal to have _1 = foo(move _1) without Miri detecting any UB. Previously this would have been caught because protecting both the destination and the moved operand would result in a conflict, which is UB. Under the new semantics, _1 is freed by operand evaluation, and then re-initialized with a new allocation by destination evaluation. This works fine in the abstract machine, but breaks in-place passing in codegen.
The solution that I chose is to make the specific case where the same local is freed and the re-initialized in the destination invalid MIR specifically for calls. But there are also alternative solutions that achieve the same effect:
- Make it UB instead of invalid MIR.
- Allow it and have codegen detect this specific situation, in which case it must treat the
move _1as a copy and not pass it in-place.
Rollup merge of #163336 - Amanieu:move-elimination/tail-copy-to-move, r=tmiasko MIR move elimination [2/6]: TailCopyToMove Depends on #163335 This PR implements a pre-pass for the `MoveElimination` pass from rust-lang/rfcs#3943. `TailCopyToMove` turns `Copy` into `Move` just before a `Return` terminator. This is valid even for places that have been borrowed because the `Return` will invalidate borrows anyway. This is necessary to allow the source of the copy to be unified with the return place. r? tmiasko
View all comments
This RFC proposes changes to Rust's operational semantics and MIR representation to enable elimination of unnecessary copies of local variables. Specifically, it makes accessing memory after a move undefined behavior, and redefines the allocation lifetime of local variables to be tied to their initialized state rather than their lexical scope. Finally, it introduces a new MIR optimization pass which exploits these guarantees to eliminate copies between locals when it is safe to do so.
Important
Since RFCs involve many conversations at once that can be difficult to follow, please use review comment threads on the text changes instead of direct comments on the RFC.
If you don't have a particular section of the RFC to comment on, you can click on the "Comment on this file" button on the top-right corner of the diff, to the right of the "Viewed" checkbox. This will create a separate thread even if others have commented on the file too.
Rendered