wasmtime

Author	SHA1	Message	Date
Nick Fitzgerald	a1f4b46f64	Bump Wasmtime to version 0.30.0; cranelift to 0.77.0	2021-09-17 10:33:50 -07:00
Nick Fitzgerald	101998733b	Merge pull request from GHSA-v4cp-h94r-m7xf Fix a use-after-free bug when passing `ExternRef`s to Wasm	2021-09-17 10:27:29 -07:00
Nick Fitzgerald	d2ce1ac753	Fix a use-after-free bug when passing `ExternRef`s to Wasm We _must not_ trigger a GC when moving refs from host code into Wasm (e.g. returned from a host function or passed as arguments to a Wasm function). After insertion into the table, this reference is no longer rooted. If multiple references are being sent from the host into Wasm and we allowed GCs during insertion, then the following events could happen: * Reference A is inserted into the activations table. This does not trigger a GC, but does fill the table to capacity. * The caller's reference to A is removed. Now the only reference to A is from the activations table. * Reference B is inserted into the activations table. Because the table is at capacity, a GC is triggered. * A is reclaimed because the only reference keeping it alive was the activation table's reference (it isn't inside any Wasm frames on the stack yet, so stack scanning and stack maps don't increment its reference count). * We transfer control to Wasm, giving it A and B. Wasm uses A. That's a use after free. To prevent uses after free, we cannot GC when moving refs into the `VMExternRefActivationsTable` because we are passing them from the host to Wasm. On the other hand, when we are cloning -- as opposed to moving -- refs from the host to Wasm, then it is fine to GC while inserting into the activations table, because the original referent that we are cloning from is still alive and rooting the ref.	2021-09-14 14:23:42 -07:00
Chris Fallin	2412e8d784	Merge pull request #3317 from dheaton-arm/implement-swiden Implement `SwidenLow` and `SwidenHigh` for the interpreter	2021-09-14 08:57:57 -07:00
dheaton-arm	99cc95d630	Factor out shared logic for widening ops. Copyright (c) 2021, Arm Limited	2021-09-14 13:08:35 +01:00
dheaton-arm	a595bd22e3	Replace loops with iterator methods. Copyright (c) 2021, Arm Limited	2021-09-14 12:37:36 +01:00
dheaton-arm	d2cbe4fc30	Fix failing test from old x86 backend Copyright (c) 2021, Arm Limited	2021-09-14 12:37:36 +01:00
dheaton-arm	75ef00f1fd	Implement `SwidenLow` and `SwidenHigh` for the interpreter Implemented `SwidenLow` and `SwidenHigh` for the Cranelift interpreter, doubling the width and halving the number of lanes preserving the low and high halves respectively. Conversions are performed using signed extension. Copyright (c) 2021, Arm Limited	2021-09-14 12:37:36 +01:00
Chris Fallin	192586506d	Merge pull request #3342 from akirilov-arm/aarch64_lowering_type_checks Cranelift AArch64: Improve the type checks for IR operations	2021-09-13 10:12:06 -07:00
Chris Fallin	7421e1a65b	Merge pull request #3324 from dheaton-arm/implement-shuffle Implement `Shuffle` for the interpreter	2021-09-13 09:49:59 -07:00
Chris Fallin	9323762d71	Merge pull request #3314 from dheaton-arm/implement-bitops Implement bit operations for Cranelift interpreter	2021-09-13 09:29:10 -07:00
Anton Kirilov	8805e25042	Cranelift AArch64: Improve the type checks for IR operations There were cases where the AArch64 backend assumed that an IR operation would always operate on certain types (the most likely reason being that the corresponding WebAssembly instruction did not cover anything else), even though the definition of the IR operation imposed no constraints like that. Copyright (c) 2021, Arm Limited.	2021-09-13 14:46:45 +01:00
Dan Gohman	256e942aa0	Tidy up redundant `use` declarations. (#3333 ) This is just a minor code cleanup.	2021-09-11 12:26:54 -05:00
Chris Fallin	587f603018	Merge pull request #3316 from dheaton-arm/implement-uwiden Implement `UwidenLow` and `UwidenHigh` for the interpreter	2021-09-10 12:32:50 -07:00
Afonso Bordado	3c1133379c	cranelift: Add `is_bool_vector` helper	2021-09-10 15:46:14 +01:00
Afonso Bordado	85d468dc5a	cranelift: Add `coerce_bools_to_ints` helper	2021-09-10 15:38:30 +01:00
Afonso Bordado	d31bdff7db	cranelift: Use bool args in simd tests	2021-09-10 15:10:51 +01:00
Afonso Bordado	9460a4fb16	cranelift: Support bool vectors in trampoline	2021-09-10 15:10:51 +01:00
dheaton-arm	4a4f940fac	Move immediate value retrieval to `imm` Copyright (c) 2021, Arm Limited	2021-09-10 12:36:33 +01:00
dheaton-arm	e7d570ddd9	Collect into Result rather than unwrap Copyright (c) 2021, Arm Limited	2021-09-10 12:26:48 +01:00
dheaton-arm	924b0368e9	Rewrite as iterator methods Copyright (c) 2021, Arm Limited	2021-09-10 09:41:23 +01:00
dheaton-arm	5824cca0f8	Fix test failures from old x86 backend Copyright (c) 2021, Arm Limited	2021-09-08 15:43:08 +01:00
dheaton-arm	f7a1b3f9bd	Implement `UwidenLow` and `UwidenHigh` for the interpreter Implemented `UwidenLow` and `UwidenHigh` for the Cranelift interpreter, doubling the width and halving the number of lanes preserving the low and high halves respectively. Conversions are performed using unsigned zero extension. Copyright (c) 2021, Arm Limited	2021-09-08 14:17:11 +01:00
dheaton-arm	dfe1c914ea	Cast types back to expected in macros Also neatened `popcnt` a little following feedback. Copyright (c) 2021, Arm Limited	2021-09-08 12:36:01 +01:00
dheaton-arm	bca3cb32ef	Implement `Shuffle` for the interpreter Implemented `Shuffle` for the Cranelift interpreter, to shuffle two SIMD vectors together based on an immediate mask of 16 bytes. Copyright (c) 2021, Arm Limited	2021-09-08 11:13:57 +01:00
dheaton-arm	9f647301ff	Implement bit operations for Cranelift interpreter Implemented for the Cranelift interpreter: - `Bitrev` to reverse the order of the bits in an integer. - `Cls` to count the leading bits which are the same as the sign bit in an integer, yielding one less than the size of the integer for 0 and -1. - `Clz` to count the number of leading zeros in the bitwise representation of the integer. - `Ctz` to count the number of trailing zeros in the bitwise representation of the integer. - `Popcnt` to count the number of ones in the bitwise representation of the integer. Copyright (c) 2021, Arm Limited	2021-09-08 11:07:22 +01:00
Afonso Bordado	3f62ef6e58	cranelift: Fix Build error #3304 and #3268 are slightly incomptible and caused the build to fail when they were merged together	2021-09-07 18:13:45 +01:00
Damian Heaton	dd23a21b9b	Implement `Swizzle` and `Splat` for interpreter (#3268 ) * Implement `Swizzle` and `Splat` for interpreter Implemented for the Cranelift interpreter: - `Swizzle` to shuffle an `i8x16` SIMD vector based on the indices specified in another vector of the same size. - `Splat` to create a SIMD vector with all lanes having the same value. Copyright (c) 2021, Arm Limited * Fix old x86 backend failing test Copyright (c) 2021, Arm Limited * Represent i16x8 and above as hex Copyright (c) 2021, Arm Limited	2021-09-07 09:53:49 -07:00
Afonso Bordado	63e9a81deb	Implement `vany_true` and `vall_true` instructions in interpreter (#3304 ) * cranelift: Implement ZeroExtend for a bunch of types in interpreter * cranelift: Implement VConst on interpreter * cranelift: Implement VallTrue on interpreter * cranelift: Implement VanyTrue on interpreter * cranelift: Mark `v{all,any}_true` tests as machinst only * cranelift: Disable `vany_true` tests on aarch64 The `b64x2` case produces an illegal instruction. See #3305	2021-09-07 09:50:39 -07:00
Afonso Bordado	81d5781e6c	cranelift: CLIF fuzzer generate jump tables and br_table	2021-09-03 19:10:49 +01:00
Afonso Bordado	cbfae6f336	cranelift: CLIF fuzzer refactor param count generation	2021-09-03 19:10:49 +01:00
Chris Fallin	36b7e81979	Merge pull request #3094 from afonso360/fuzzer-branching Cranelift CLIF Fuzzer generate blocks and branches	2021-09-03 10:30:14 -07:00
Chris Fallin	ecd795f736	Merge pull request #3290 from dheaton-arm/implement-ssatarith Implement `SaddSat` and `SsubSat` for the Cranelift interpreter	2021-09-03 09:48:34 -07:00
Chris Fallin	e3ccff0249	Merge pull request #3283 from dheaton-arm/implement-umulhi Implement `Umulhi` for the interpreter	2021-09-03 09:29:21 -07:00
dheaton-arm	8f057e0482	Implement `SaddSat` and `SsubSat` for the interpreter Implemented `SaddSat` and `SsubSat` to add and subtract signed vector values, saturating at the type boundaries rather than overflowing. Changed the parser to allow signed `i8` immediates in vectors as part of this work; fixes #3276. Copyright (c) 2021, Arm Limited.	2021-09-03 11:35:39 +01:00
dheaton-arm	562947c678	Fix CI tests + rename tests - Fixed CI tests for AArch64 and old x86. - Rename `simd-umulhi.clif` to `umulhi.clif`. - Rename `simd-umulhi-aarch64.clif` to `simd-umulhi.clif`. Copyright (c) 2021, Arm Limited.	2021-09-03 10:37:24 +01:00
Chris Fallin	d6a77898ba	Merge pull request #3272 from dheaton-arm/implement-iaddpairwise Implement `IaddPairwise` for the interpreter	2021-09-02 10:52:47 -07:00
Chris Fallin	6e05b646a3	Merge pull request #3282 from afonso360/x64-fix-brtables cranelift: Fix `br_table` for `i64` types in x64 backend.	2021-09-02 09:58:42 -07:00
Chris Fallin	2389a4ea00	Merge pull request #3274 from bnjbvr/fix-m1 A round of Mac M1 fixes	2021-09-02 09:49:01 -07:00
Chris Fallin	000a97f4ff	Merge pull request #3279 from dheaton-arm/implement-insertlane Implement `Insertlane` for the Cranelift interpreter	2021-09-02 09:44:59 -07:00
Afonso Bordado	20913a7473	cranelift: Enable compiling br_tables for types larger than i32	2021-09-02 16:26:23 +01:00
Afonso Bordado	f9ada24bcf	cranelift: Fix br_table for i64 inputs We still only support a maximum of u32::MAX entries, however we no longer crash when compiling 64 bit indexes. Fixes #3100	2021-09-02 15:31:48 +01:00
dheaton-arm	16b6a404e4	Implement `Umulhi` for the interpreter Implemented `Umulhi` for the Cranelift interpreter, performing unsigned integer multiplication and producing the high half of a double-length result. Fixed `ExtractUpper` conversion behaviour as part of this change, which was extracting from a 128-bit value regardless of the size of the original value. Copyright (c) 2021, Arm Limited.	2021-09-02 13:11:41 +01:00
Benjamin Bouvier	85ec11acb9	Aarch64: always generate the CFA directive indicating no pointer signing	2021-09-02 09:16:34 +02:00
Benjamin Bouvier	fb94b81538	Use 16K code pages on Mac M1 Fixes #3278.	2021-09-02 09:16:34 +02:00
Afonso Bordado	f0f2efba26	cranelift: CLIF Fuzzer generate brz/brnz/bricmp instructions	2021-09-01 19:44:11 +01:00
Afonso Bordado	f4bd7d17a3	cranelift: CLIF Fuzzer generate multiple blocks	2021-09-01 19:44:05 +01:00
Alex Crichton	1532516a36	Use relative `call` instructions between wasm functions (#3275 ) * Use relative `call` instructions between wasm functions This commit is a relatively major change to the way that Wasmtime generates code for Wasm modules and how functions call each other. Prior to this commit all function calls between functions, even if they were defined in the same module, were done indirectly through a register. To implement this the backend would emit an absolute 8-byte relocation near all function calls, load that address into a register, and then call it. While this technique is simple to implement and easy to get right, it has two primary downsides associated with it: * Function calls are always indirect which means they are more difficult to predict, resulting in worse performance. * Generating a relocation-per-function call requires expensive relocation resolution at module-load time, which can be a large contributing factor to how long it takes to load a precompiled module. To fix these issues, while also somewhat compromising on the previously simple implementation technique, this commit switches wasm calls within a module to using the `colocated` flag enabled in Cranelift-speak, which basically means that a relative call instruction is used with a relocation that's resolved relative to the pc of the call instruction itself. When switching the `colocated` flag to `true` this commit is also then able to move much of the relocation resolution from `wasmtime_jit::link` into `wasmtime_cranelift::obj` during object-construction time. This frontloads all relocation work which means that there's actually no relocations related to function calls in the final image, solving both of our points above. The main gotcha in implementing this technique is that there are hardware limitations to relative function calls which mean we can't simply blindly use them. AArch64, for example, can only go +/- 64 MB from the `bl` instruction to the target, which means that if the function we're calling is a greater distance away then we would fail to resolve that relocation. On x86_64 the limits are +/- 2GB which are much larger, but theoretically still feasible to hit. Consequently the main increase in implementation complexity is fixing this issue. This issue is actually already present in Cranelift itself, and is internally one of the invariants handled by the `MachBuffer` type. When generating a function relative jumps between basic blocks have similar restrictions. This commit adds new methods for the `MachBackend` trait and updates the implementation of `MachBuffer` to account for all these new branches. Specifically the changes to `MachBuffer` are: * For AAarch64 the `LabelUse::Branch26` value now supports veneers, and AArch64 calls use this to resolve relocations. * The `emit_island` function has been rewritten internally to handle some cases which previously didn't come up before, such as: * When emitting an island the deadline is now recalculated, where previously it was always set to infinitely in the future. This was ok prior since only a `Branch19` supported veneers and once it was promoted no veneers were supported, so without multiple layers of promotion the lack of a new deadline was ok. * When emitting an island all pending fixups had veneers forced if their branch target wasn't known yet. This was generally ok for 19-bit fixups since the only kind getting a veneer was a 19-bit fixup, but with mixed kinds it's a bit odd to force veneers for a 26-bit fixup just because a nearby 19-bit fixup needed a veneer. Instead fixups are now re-enqueued unless they're known to be out-of-bounds. This may run the risk of generating more islands for 19-bit branches but it should also reduce the number of islands for between-function calls. * Otherwise the internal logic was tweaked to ideally be a bit more simple, but that's a pretty subjective criteria in compilers... I've added some simple testing of this for now. A synthetic compiler option was create to simply add padded 0s between functions and test cases implement various forms of calls that at least need veneers. A test is also included for x86_64, but it is unfortunately pretty slow because it requires generating 2GB of output. I'm hoping for now it's not too bad, but we can disable the test if it's prohibitive and otherwise just comment the necessary portions to be sure to run the ignored test if these parts of the code have changed. The final end-result of this commit is that for a large module I'm working with the number of relocations dropped to zero, meaning that nothing actually needs to be done to the text section when it's loaded into memory (yay!). I haven't run final benchmarks yet but this is the last remaining source of significant slowdown when loading modules, after I land a number of other PRs both active and ones that I only have locally for now. * Fix arm32 * Review comments	2021-09-01 13:27:38 -05:00
Chris Fallin	91410aaddf	Merge pull request #3234 from dheaton-arm/implement-isubb Implement `IsubBin`, `IsubBout`, and `IsubBorrow`for Cranelift interpreter	2021-09-01 11:25:43 -07:00
Chris Fallin	2a63979151	Merge pull request #3258 from afonso360/ssa-dce cranelift: Prevent infinite loops in ssa frontend with unreachable code.	2021-09-01 11:19:40 -07:00

1 2 3 4 5 ...

3181 Commits