wasmtime

Author	SHA1	Message	Date
Alex Crichton	2d25db047f	x64: Lower SIMD requirement to SSE4.1 from SSE4.2 (#6206 ) Cranelift only has one instruction SIMD which depends on SSE4.2 so this commit adds a lowering rule for `pcmpgtq` which doesn't use SSE4.2 and enables lowering the baseline requirement for SIMD support from SSE4.2 to SSE4.1. The `has_sse42` setting is no longer enabled by default for Cranelift. Additionally `enable_simd` no longer requires `has_sse42` on x64. Finally the fuzz-generator for Wasmtime codegen settings now enables flipping the `has_sse42` setting instead of unconditionally setting it to `true`. The specific lowering for `pcmpgtq` is copied from LLVM's lowering of this instruction.	2023-04-14 17:24:43 +00:00
T0b1-iOS	3956a6aa0f	remove `unsigned_add_overflow_condition` (#6199 )	2023-04-13 14:30:44 +00:00
Karl Meakin	91e36f3449	Clarify the representation of `icmp` output (#6202 ) * Clarify the representation of `icmp` output * Reformat * "ie" => "i.e." * Update `fcmp` documentation as well	2023-04-12 20:05:44 +00:00
Karl Meakin	42528d82b8	Add `multi_lane` precondition to `bitselect` => `{u,s}{min,max}` rewrite (#6201 )	2023-04-12 19:04:30 +00:00
T0b1-iOS	f684a5fbee	remove `iadd_cout` and `isub_bout` (#6198 )	2023-04-11 23:39:32 +00:00
Karl Meakin	c0166f78f9	ISLE: simplify select/bitselect when both choices are the same (#6141 )	2023-04-11 22:41:19 +00:00
Karl Meakin	b9a58148cf	ISLE: split algebraic.isle into several files (#6140 ) * ISLE: split algebraic.isle into several files * delete `algebraic.clif` * Add `README.md` * Remove old `algebraic.clif` tests --------- Co-authored-by: Jamey Sharp <jsharp@fastly.com>	2023-04-11 21:39:18 +00:00
T0b1-iOS	569089e473	Add `{u,s}{add,sub,mul}_overflow` instructions (#5784 ) * add `{u,s}{add,sub,mul}_overflow` with interpreter * add `{u,s}{add,sub,mul}_overflow` for x64 * add `{u,s}{add,sub,mul}_overflow` for aarch64 * 128bit filetests for `{u,s}{add,sub,mul}_overflow` * `{u,s}{add,sub,mul}_overflow` emit tests for x64 * `{u,s}{add,sub,mul}_overflow` emit tests for aarch64 * Initial review changes * add `with_flags_extended` helper * add `with_flags_chained` helper	2023-04-11 20:16:04 +00:00
Afonso Bordado	4c32dd7786	riscv64: Delete `SelectIf` instruction (#5888 ) * riscv64: Delete `SelectIf` instruction * riscv64: Fix typo in comment Co-authored-by: Trevor Elliott <awesomelyawesome@gmail.com> * riscv64: Improve `bmask` codegen * riscv64: Use `lower_bmask` in `select_spectre_guard` * riscv64: Use `lower_bmask` to extend values in `select_spectre_guard` Co-authored-by: Trevor Elliott <awesomelyawesome@gmail.com> --------- Co-authored-by: Trevor Elliott <awesomelyawesome@gmail.com>	2023-04-11 17:33:32 +00:00
kevaundray	f2393b8f27	Removes debug assertion that was related to issue 796 (#6175 ) * fix typo: behaviour -> behavior * remove debug assertion since 796 has been merged * Update data_value.rs	2023-04-11 16:52:11 +00:00
bjorn3	0478ead3f8	Handle signature() for more libcalls (#6174 ) * Handle signature() for more libcalls This is necessary to be able to call them in the interpreter. All the remaining libcalls which signature() doesn't handle are never used in clif ir. Only in code compiled by a backend. * Fix libcall declarations in cranelift-frontend * Add function signatures * Use correct pointer type instead of I64	2023-04-11 16:50:41 +00:00
kevaundray	4053ae9e08	Minir typo/Grammar fixes (#6187 ) * fix typo * add test to check that Option<EntityRef> is twice as large as EntityRef * grammar * grammar * reverse snakecase -- Not sure if folks want this type of change	2023-04-10 19:39:25 +00:00
Alex Crichton	435b6894d7	x64: Clarify and shrink up ModRM/SIB encoding (#6181 ) I noticed recently that for the `ImmRegRegShift` addressing mode Cranelift will unconditionally emit at least a 1-byte immediate for the offset to be added to the register addition computation, even when the offset is zero. In this case though the instruction encoding can be slightly more compact and remove a byte. This commit started off by applying this optimization, which resulted in the `*.clif` test changes in this commit. Further reading this code, however, I personally found it quite hard to follow what was happening with all the various branches and ModRM/SIB bits. I reviewed these encodings in the x64 architecture manual and attempted to improve the logic for encoding here. The new version in this commit is intended to be functionally equivalent to the prior version where dropping a zero-offset from the `ImmRegRegShift` variant is the only change.	2023-04-10 19:37:19 +00:00
Chris Fallin	8f1a7773a3	Revert "ISLE: rewrite loose inequalities to strict inequalities and strict inequalities to equalities (#6130 )" (#6193 ) This reverts commit `57e42d0c46`. Fixes #6185.	2023-04-10 18:43:15 +00:00
bjorn3	b9fb31e9a7	Re-export cranelift-control from cranelift-codegen (#6173 ) This makes it easier to keep the versions of both in sync and avoids having to specify another dependency for a single type.	2023-04-10 16:49:43 +00:00
Chris Dickinson	a97e82c6e2	doc: fix StackSlot reference to FunctionBuilder (#6182 ) `FunctionBuilder::create_stackslot` was split into `create_sized_stack_slot` and `create_dynamic_stack_slot`. This updates the doc in the `StackBuilder` docstring to refer to the new methods. Fixes #5838.	2023-04-09 21:14:19 +00:00
Alexa VanHattum	71d3b638f3	Clarify instructions.rs documentation for ushr/ashr (narrow values) (#6186 )	2023-04-09 20:01:49 +00:00
kevaundray	e3dbad9cc2	add result type assertion (#6184 )	2023-04-09 19:55:15 +00:00
Chris Fallin	230e2135d6	Cranelift: remove non-egraphs optimization pipeline and `use_egraphs` option. (#6167 ) * Cranelift: remove non-egraphs optimization pipeline and `use_egraphs` option. This PR removes the LICM, GVN, and preopt passes, and associated support pieces, from `cranelift-codegen`. Not to worry, we still have optimizations: the egraph framework subsumes all of these, and has been on by default since #5181. A few decision points: - Filetests for the legacy LICM, GVN and simple_preopt were removed too. As we built optimizations in the egraph framework we wrote new tests for the equivalent functionality, and many of the old tests were testing specific behaviors in the old implementations that may not be relevant anymore. However if folks prefer I could take a different approach here and try to port over all of the tests. - The corresponding filetest modes (commands) were deleted too. The `test alias_analysis` mode remains, but no longer invokes a separate GVN first (since there is no separate GVN that will not also do alias analysis) so the tests were tweaked slightly to work with that. The egrpah testsuite also covers alias analysis. - The `divconst_magic_numbers` module is removed since it's unused without `simple_preopt`, though this is the one remaining optimization we still need to build in the egraphs framework, pending #5908. The magic numbers will live forever in git history so removing this in the meantime is not a major issue IMHO. - The `use_egraphs` setting itself was removed at both the Cranelift and Wasmtime levels. It has been marked deprecated for a few releases now (Wasmtime 6.0, 7.0, upcoming 8.0, and corresponding Cranelift versions) so I think this is probably OK. As an alternative if anyone feels strongly, we could leave the setting and make it a no-op. * Update test outputs for remaining test differences.	2023-04-06 18:11:03 +00:00
Afonso Bordado	a9cda5af19	cranelift: Implement PartialEq in `Function` (#6157 )	2023-04-05 22:33:10 +00:00
Remo Senekowitsch	7eb8914090	Chaos mode MVP: Skip branch optimization in MachBuffer (#6039 ) * fuzz: Add chaos mode control plane Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * fuzz: Skip branch optimization with chaos mode Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * fuzz: Rename chaos engine -> control plane Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * chaos mode: refactoring ControlPlane to be passed through the call stack by reference Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Remo Senekowitsch <contact@remsle.dev> * fuzz: annotate chaos todos Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * fuzz: cleanup control plane Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * fuzz: remove control plane from compiler context Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * fuzz: move control plane into emit state Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * fuzz: fix remaining compiler errors Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * fix tests * refactor emission state ctrl plane accessors Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * centralize conditional compilation of chaos mode Also cleanup a few straggling dependencies on cranelift-control that aren't needed anymore. Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * add cranelift-control to published crates prtest:full Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> * add cranelift-control to public crates Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> --------- Co-authored-by: Falk Zwimpfer <24669719+FalkZ@users.noreply.github.com> Co-authored-by: Moritz Waser <mzrw.dev@pm.me> Co-authored-by: Remo Senekowitsch <contact@remsle.dev>	2023-04-05 19:28:46 +00:00
Alex Crichton	967543eb43	aarch64: Add more lowerings for the CLIF `fma` (#6150 ) This commit adds new lowerings to the AArch64 backend of the element-based `fmla` and `fmls` instructions. These instructions have one of the multiplicands as an implicit broadcast of a single lane of another register and can help remove `shuffle` or `dup` instructions that would otherwise be used to implement them.	2023-04-05 17:22:55 +00:00
wasmtime-publish	bf741955f0	Bump Wasmtime to 9.0.0 (#6143 ) Co-authored-by: Wasmtime Publish <wasmtime-publish@users.noreply.github.com>	2023-04-05 17:06:36 +00:00
Alex Crichton	d45cbba83f	Add egraph cprop optimizations for `splat` (#6148 ) This commit adds constant-propagation optimizations for `splat`-of-constant to produce a `vconst` node. This should help later hoisting these constants out of loops if it shows up in wasm.	2023-04-05 16:10:45 +00:00
Jamey Sharp	81545c3a86	Revert "simple_gvn: recognize commutative operators (#6135 )" (#6142 ) This reverts commit `c85bf27ff8`.	2023-04-04 20:22:44 +00:00
Karl Meakin	57e42d0c46	ISLE: rewrite loose inequalities to strict inequalities and strict inequalities to equalities (#6130 ) * ISLE: rewrite loose inequalities to strict inequalities * Rewrite strict inequalities to equalities where possible	2023-04-04 17:42:19 +00:00
Karl Meakin	c85bf27ff8	simple_gvn: recognize commutative operators (#6135 ) * simple_gvn: recognize commutative operators Normalize instructions with commutative opcodes by sorting the arguments. This means instructions like `iadd v0, v1` and `iadd v1, v0` will be considered identical by GVN and deduplicated. * Remove `UsubSat` and `SsubSat` from `is_commutative` They are not actually commutative * Remove `TODO`s * Move InstructionData normalization into helper fn * Add normalization of commutative instructions in the epgrah implementation * Handle reflexive icmp/fcmps in GVN * Change formatting of `normalize_in_place` * suggestions from code review	2023-04-04 00:25:05 +00:00
Karl Meakin	c8c224ead6	ISLE: move `icmp` rewrites to separate file. (#6120 ) * ISLE: move `icmp` rewrites to separate file. Move `icmp`-related rewrite rules from `algebraic.isle` to `icmp.isle`. Also move `icmp`-related tests from `algebraic.clif` to `icmp.clif`. * Put parameterized and unparameterized `icmp` tests in separate files * Undo refactoring of (ir)reflexivity rewrites * Fix `icmp-parameterised.clif` * Undo formatting/comment changes	2023-03-31 17:40:31 +00:00
Yoni L	94f2ff0921	cranelift::codegen::Context::optimize(): reduce verbosity of "egraph stats" traces (#6122 )	2023-03-30 00:46:14 +00:00
Alex Crichton	0b0ac3ff73	x64: Add AVX support for some more float-related instructions (#6092 ) * x64: Add AVX encodings of `vcvt{ss2sd,sd2ss}` Additionally update the instruction helpers to take an `XmmMem` argument to allow load sinking into the instruction. * x64: Add AVX encoding of `sqrts{s,d}` * x64: Add AVX support for `rounds{s,d}`	2023-03-29 18:09:49 +00:00
Alex Crichton	afb417920d	x64: Deduplicate fcmp emission logic (#6113 ) * x64: Deduplicate fcmp emission logic The `select`-of-`fcmp` lowering duplicated a good deal of `FloatCC` lowering logic that was already done by `emit_fcmp`, so this commit refactors these lowering rules to instead delegate to `emit_fcmp` and then handle that result. * Swap order of condition codes Shouldn't affect the correctness of this operation and it's a bit more natural to write the lowering rule this way. * Swap the order of comparison operands No need to swap `a b`, only the `x y` needs swapping. * Fix x64 printing of `XmmCmove`	2023-03-29 16:24:25 +00:00
Karl Meakin	dcf0ea9ff3	ISLE: rewrite `and`/`or` of `icmp` (#6095 ) * ISLE: rewrite `and`/`or` of `icmp` * Add `make-icmp-tests.sh` script * Remove unused changes	2023-03-29 00:13:27 +00:00
Maja Kądziołka	db07988ccb	x64: emit_cmp: use x64_test for comparisons with 0 (#6086 ) * x64: emit_cmp: use x64_test for comparisons with 0 See #5869 * fixup! x64: emit_cmp: use x64_test for comparisons with 0	2023-03-27 15:38:48 +00:00
Afonso Bordado	a002a2cc5e	riscv64: Add instruction helpers (#6099 ) * riscv64: Add helpers for `add` * riscv64: Add helpers for `sub` * riscv64: Add helpers for `sll` * riscv64: Add helpers for `srl` * riscv64: Add helpers for `sra` * riscv64: Add helpers for `or` * riscv64: Add helpers for `and` * riscv64: Add helpers for `xor` * riscv64: Add helpers for `addi` * riscv64: Add helpers for `slli` * riscv64: Add helpers for `srli` * riscv64: Add helpers for `srai` * riscv64: Add helpers for `ori` * riscv64: Add helpers for `xori` * riscv64: Add helpers for `andi` * riscv64: Add helpers for `not` * riscv64: Add helpers for `sltiu` * riscv64: Add helpers for `seqz` * riscv64: Add helpers for `addw` * riscv64: Add helpers for `subw` * riscv64: Add helpers for `sllw` * riscv64: Add helpers for `slliw` * riscv64: Add helpers for `srlw` * riscv64: Add helpers for `srliw` * riscv64: Add helpers for `sraw` * riscv64: Add helpers for `sraiw` * riscv64: Add helpers for `sltu` * riscv64: Add helpers for `mul` * riscv64: Add helpers for `mulh` * riscv64: Add helpers for `mulhu` * riscv64: Add helpers for `div` * riscv64: Add helpers for `divu` * riscv64: Add helpers for `rem` * riscv64: Add helpers for `remu` * riscv64: Add helpers for `mulw` * riscv64: Add helpers for `divw` * riscv64: Add helpers for `divuw` * riscv64: Add helpers for `remw` * riscv64: Add helpers for `remuw` * riscv64: Add helpers for `neg` * riscv64: Add helpers for `addiw` * riscv64: Add helpers for `sext.w` * riscv64: Add helpers for `fadd` * riscv64: Add helpers for `fsub` * riscv64: Add helpers for `fmul` * riscv64: Add helpers for `fdiv` * riscv64: Add helpers for `fsqrt` * riscv64: Add helpers for `fmadd` * riscv64: Add helpers for `fsgnj` * riscv64: Add helpers for `fsgnjn` * riscv64: Add helpers for `fsgnjx` * riscv64: Add helpers for `fcvtds` * riscv64: Add helpers for `fcvtsd` * riscv64: Add helpers for `adduw` * riscv64: Add helpers for `zext.w` * riscv64: Add helpers for `andn` * riscv64: Add helpers for `orn` * riscv64: Add helpers for `clz` * riscv64: Add helpers for `clzw` * riscv64: Add helpers for `ctz` * riscv64: Add helpers for `ctzw` * riscv64: Add helpers for `cpop` * riscv64: Add helpers for `max` * riscv64: Add helpers for `feq` * riscv64: Add helpers for `flt` * riscv64: Add helpers for `fle` * riscv64: Add helpers for `fgt` * riscv64: Add helpers for `fge` * riscv64: Add helpers for `sext.b` * riscv64: Add helpers for `sext.h` * riscv64: Add helpers for `zext.h` * riscv64: Add helpers for `rol` * riscv64: Add helpers for `rolw` * riscv64: Add helpers for `ror` * riscv64: Add helpers for `rorw` * riscv64: Add helpers for `rev8` * riscv64: Add helpers for `brev8` * riscv64: Add helpers for `bseti` * riscv64: Add helpers for `pack` * riscv64: Add helpers for `packw` * riscv64: Add helpers for `slli.uw` * riscv64: Add helpers for `fabs` * riscv64: Add helpers for `fneg`	2023-03-24 18:01:04 +00:00
Nathan Whitaker	c3decdf910	cranelift: Implement TLS on aarch64 Mach-O (Apple Silicon) (#5434 ) * Implement TLS on Aarch64 Mach-O * Add aarch64 macho TLS filetest * Address review comments - `Aarch64` instead of `AArch64` in comments - Remove unnecessary guard in tls_value lowering - Remove unnecessary regalloc metadata in emission * Use x1 as temporary register in emission - Instead of passing in a temporary register to use when emitting the TLS code, just use `x1`, as it's already in the clobber set. This also keeps the size of `aarch64::inst::Inst` at 32 bytes. - Update filetest accordingly * Update aarch64 mach-o TLS filetest	2023-03-24 17:54:01 +00:00
Afonso Bordado	3546ccf7d1	riscv64: Cleanup unused `lower_float_unordered` (#6096 )	2023-03-23 21:08:38 +00:00
Afonso Bordado	602ff71fe4	riscv64: Add `Zba` extension instructions (#6087 ) * riscv64: Use `add.uw` to zero extend * riscv64: Implement `add.uw` optimizations * riscv64: Add `Zba` `iadd+ishl` optimizations * riscv64: Add `shl+uextend` optimizations based on `Zba` * riscv64: Fix some issues with `Zba` instructions * riscv64: Restrict shnadd selection * riscv64: Fix `extend` priorities * riscv64: Remove redundant `addw` rule * riscv64: Specify type for `add` extend rules * riscv64: Use `u64_from_imm64` extractor instead of `uimm8` * riscv64: Restrict `uextend` in `shnadd.uw` rules * riscv64: Use concrete type in `slli.uw` rule * riscv64: Add extra arithmetic extends tests Co-authored-by: Jamey Sharp <jsharp@fastly.com> * riscv64: Make `Adduw` types concrete * riscv64: Add extra arithmetic extend tests * riscv64: Add `sextend`+Arithmetic rules * riscv64: Fix whitespace * cranelift: Move arithmetic extends tests with i128 to separate file --------- Co-authored-by: Jamey Sharp <jsharp@fastly.com>	2023-03-23 20:06:03 +00:00
Ulrich Weigand	6f66abd5c7	s390x: Improved TrapIf implementation (#6079 ) Following up on the discussion in https://github.com/bytecodealliance/wasmtime/pull/6011 this adds an improved implementation of TrapIf for s390x using a single conditional branch instruction. If the trap conditions is true, we branch into the middle of the branch instruction - those middle two bytes are zero, which matches the encoding of the trap instruction. In addition, show the trap code for Trap and TrapIf instructions in assembler output.	2023-03-23 14:50:43 +00:00
Alex Crichton	2fde25311e	x64: Refactor and fill out some gpr-vs-xmm bits (#6058 ) * x64: Add instruction helpers for `mov{d,q}` These will soon grow AVX-equivalents so move them to instruction helpers to have clauses for AVX in the future. * x64: Don't auto-convert between RegMemImm and XmmMemImm The previous conversion, `mov_rmi_to_xmm`, would move from GPR registers to XMM registers which isn't what many of the other `convert` statements between these newtypes do. This seemed like a possible footgun so I've removed the auto-conversion and added an explicit helper to go from a `u32` to an `XmmMemImm`. * x64: Add AVX encodings of some more GPR-related insns This commit adds some more support for AVX instructions where GPRs are in use mixed in with XMM registers. This required a few more variants of `Inst` to handle the new instructions. * Fix vpmovmskb encoding * Fix xmm-to-gpr encoding of vmovd/vmovq * Fix typo * Fix rebase conflict * Fix rebase conflict with tests	2023-03-22 14:58:09 +00:00
Afonso Bordado	7a3df7dcc0	riscv64: Improve `ctz`/`clz`/`cls` codegen (#5854 ) * cranelift: Add extra runtests for `clz`/`ctz` * riscv64: Restrict lowering rules for `ctz`/`clz` * cranelift: Add `u64` isle helpers * riscv64: Improve `ctz` codegen * riscv64: Improve `clz` codegen * riscv64: Improve `cls` codegen * riscv64: Improve `clz.i128` codegen Instead of checking if we have 64 zeros in the top half. Check if it is 0, that way we avoid loading the `64` constant. * riscv64: Improve `ctz.i128` codegen Instead of checking if we have 64 zeros in the bottom half. Check if it is 0, that way we avoid loading the `64` constant. * riscv64: Use extended value in `lower_cls` * riscv64: Use pattern matches on `bseti`	2023-03-21 23:15:14 +00:00
Karl Meakin	ff6f17ca52	ISLE: add synonyms for all variations of `icmp` (#6081 )	2023-03-21 22:13:00 +00:00
Alexa VanHattum	13be5618a7	Cranelift: ISLE: aarch64: fix `imm12_from_negated_value` for `i32`, `i16` (#6078 ) * Fix the semantics of imm12_from_negated_value, swapping to a partial term + rule * wrapping_neg	2023-03-21 19:16:25 +00:00
Trevor Elliott	861220c433	Restrict the types for isplit and iconcat to match backends (#6070 ) * Restrict the types for isplit and iconcat to match backends * Admit unimplemented bitwidths to isplit/iconcat * Modify the NarrowInt type instead of shadowing it * Fix filetest failures	2023-03-21 01:21:00 +00:00
Karl Meakin	7d9318fe77	cranelift: rewrite `iabs(ineg(x))` and `iabs(iabs(x))` (#6072 ) * cranelift: rerwite `iabs(ineg(x))`` and `iabs(iabs(x))` * Fix comment on `iabs(iabs(x))` rewrite * Remove subsume on rewrite for `iabs(ineg(x))`	2023-03-21 00:12:21 +00:00
Alex Crichton	a3b21031d4	Add a `MachBuffer::defer_trap` method (#6011 ) * Add a `MachBuffer::defer_trap` method This commit adds a new method to `MachBuffer` to defer trap opcodes to the end of a function in a similar manner to how constants are deferred to the end of the function. This is useful for backends which frequently use `TrapIf`-style opcodes. Currently a jump is emitted which skips the next instruction, a trap, and then execution continues normally. While there isn't any pressing problem with this construction the trap opcode is in the middle of the instruction stream as opposed to "off on the side" despite rarely being taken. With this method in place all the backends (except riscv64 since I couldn't figure it out easily enough) have a new lowering of their `TrapIf` opcode. Now a trap is deferred, which returns a label, and then that label is jumped to when executing the trap. A fixup is then recorded in `MachBuffer` to get patched later on during emission, or at the end of the function. Subsequently all `TrapIf` instructions translate to a single branch plus a single trap at the end of the function. I've additionally further updated some more lowerings in the x64 backend which were explicitly using traps to instead use `TrapIf` where applicable to avoid jumping over traps mid-function. Other backends didn't appear to have many jump-over-the-next-trap patterns. Lots of tests have had their expectations updated here which should reflect all the traps being sunk to the end of functions. * Print trap code on all platforms * Emit traps before constants * Preserve source location information for traps * Fix test expectations * Attempt to fix s390x The MachBuffer was registering trap codes with the first byte of the trap, but the SIGILL handler was expecting it to be registered with the last byte of the trap. Exploit that SIGILL is always represented with a 2-byte instruction and always march 2-backwards for SIGILL, continuing to march backwards 1 byte for SIGFPE-generating instructions. * Back out s390x changes * Back out more s390x bits * Review comments	2023-03-20 21:24:47 +00:00
bjorn3	49bab6db7f	Ensure the sequence number doesn't leak out of Layout (#6061 ) Previously it could affect the PartialEq and Hash impls. Ignoring the sequence number in PartialEq and Hash allows us to not renumber all blocks in the incremental cache.	2023-03-20 19:20:00 +00:00
bjorn3	fc3c5d2414	Properly use the VersionMarker in CachedFunc (#6062 )	2023-03-20 19:18:51 +00:00
Alex Crichton	f7dda1ab2c	x64: Fix vbroadcastss with AVX2 and without AVX (#6060 ) * x64: Fix vbroadcastss with AVX2 and without AVX This commit fixes a corner case in the emission of the `vbroadcasts{s,d}` instructions. The memory-to-xmm form of these instructions was available with the AVX instruction set, but the xmm-to-xmm form of these instructions wasn't available until AVX2. The instruction requirement for these are listed as AVX but the lowering rules are appropriately annotated to use either AVX2 or AVX when appropriate. While this should work in practice this didn't work for the assertion about enabled features for each instruction. The `vbroadcastss` instruction was listed as requiring AVX but could get emitted when AVX2 was enabled (due to the reg-to-reg form being available). This caused an issue for the fuzzer where AVX2 was enabled but AVX was disabled. One possible fix would be to add more opcodes, one for reg-to-reg and one for mem-to-reg. That seemed like somewhat overkill for a pretty niche situation that shouldn't actually come up in practice anywhere. Instead this commit changes all the `has_avx` accessors to the `use_avx_simd` predicate already available in the target flags. The `use_avx2_simd` predicate was then updated to additionally require `has_avx`, so if AVX2 is enabled and AVX is disabled then the `vbroadcastss` instruction won't get emitted any more. Closes #6059 * Pass `enable_simd` on a few more files	2023-03-18 18:38:03 +00:00
Trevor Elliott	78dbe93f21	Rename `as_bool` to `as_truthy`, and fix TypeSet::as_bool (#6027 )	2023-03-17 21:11:24 +00:00
bjorn3	2c40c267d4	Make sequence numbers local to instructions (#6043 ) * Only allow pp_cmp within a single block Block order shouldn't matter for codegen and restricting pp_cmp to a single block will allow making instruction sequence numbers local to a block. * Make sequence numbers local to instructions This allows renumbering to be localized to a single block where previously it could affect the entire function. Also saves 32bit of overhead per block.	2023-03-17 20:53:21 +00:00

1 2 3 4 5 ...

2318 Commits