wasmtime

Author	SHA1	Message	Date
bjorn3	53ec12d519	Rustfmt	2021-09-29 16:27:47 +02:00
bjorn3	9e5201d88f	Fix all dead-code warnings in cranelift-codegen	2021-09-29 16:27:47 +02:00
bjorn3	3e4167ba95	Remove registers from cranelift-codegen-meta	2021-09-29 16:27:47 +02:00
bjorn3	18bd27e90b	Remove legalizer support from cranelift-codegen-meta	2021-09-29 16:27:45 +02:00
bjorn3	d499933612	Remove encoding generation from cranelift-codegen-meta	2021-09-29 16:23:58 +02:00
bjorn3	d8818c967e	Fix all dead-code warnings in cranelift-codegen-meta	2021-09-29 16:23:58 +02:00
bjorn3	59e18b7d1b	Remove the old riscv backend	2021-09-29 16:23:57 +02:00
bjorn3	9e34df33b9	Remove the old x86 backend	2021-09-29 16:13:46 +02:00
Alex Crichton	1ee2af0098	Remove the lightbeam backend (#3390 ) This commit removes the Lightbeam backend from Wasmtime as per [RFC 14]. This backend hasn't received maintenance in quite some time, and as [RFC 14] indicates this doesn't meet the threshold for keeping the code in-tree, so this commit removes it. A fast "baseline" compiler may still be added in the future. The addition of such a backend should be in line with [RFC 14], though, with the principles we now have for stable releases of Wasmtime. I'll close out Lightbeam-related issues once this is merged. [RFC 14]: https://github.com/bytecodealliance/rfcs/pull/14	2021-09-27 12:27:19 -05:00
Chris Fallin	344a219245	Merge pull request #3383 from akirilov-arm/vany_true Cranelift AArch64: Fix the VanyTrue implementation for 64-bit elements	2021-09-24 09:26:36 -07:00
Anton Kirilov	0fb3acfb94	Cranelift AArch64: Fix the VanyTrue implementation for 64-bit elements Copyright (c) 2021, Arm Limited.	2021-09-23 20:39:46 +01:00
Anton Kirilov	930b1f17f0	Cranelift AArch64: Implement scalar FmaxPseudo and FminPseudo Copyright (c) 2021, Arm Limited.	2021-09-23 15:11:01 +01:00
Chris Fallin	3474965ca6	Merge pull request #3322 from sparker-arm/aarch64-lse-ops AArch64 LSE atomic_rmw support	2021-09-22 09:21:28 -07:00
Chris Fallin	38728c5746	Merge pull request #3362 from dheaton-arm/implement-unarrow Implement `Unarrow`, `Uunarrow`, and `Snarrow` for the interpreter	2021-09-21 10:06:46 -07:00
Chris Fallin	ebe2af6eaa	Merge pull request #3351 from afonso360/parser-i128 cranelift: Add support for parsing i128 data values	2021-09-21 10:04:27 -07:00
Ulrich Weigand	51131a3acc	Fix s390x regressions (#3330 ) - Add relocation handling needed after PR #3275 - Fix incorrect handling of signed constants detected by PR #3056 test - Fix LabelUse max pos/neg ranges; fix overflow in buffers.rs - Disable fuzzing tests that require pre-built v8 binaries - Disable cranelift test that depends on i128 - Temporarily disable memory64 tests	2021-09-20 09:12:36 -05:00
dheaton-arm	3fc29f5f6c	Return `u128` from `bounds`; form `new_vec` from iter chain Copyright (c) 2021, Arm Limited	2021-09-20 09:57:19 +01:00
Afonso Bordado	8115e7252d	cranelift: Add support for i128 immediates in parser	2021-09-19 15:02:04 +01:00
Chris Fallin	6a98fe2104	Merge pull request #3332 from afonso360/interp-icmp cranelift: Add SIMD `icmp` to interpreter	2021-09-17 15:13:44 -07:00
Chris Fallin	c9834ee91c	Merge pull request #3329 from uweigand/datavalue-endian-fix cranelift: Fix big-endian regression in data_value.rs	2021-09-17 14:55:07 -07:00
Chris Fallin	1f2d1c097d	Merge pull request #3364 from dheaton-arm/implement-smulhi Implement `Smulhi` for interpreter	2021-09-17 12:56:37 -07:00
Chris Fallin	94151434ce	Merge pull request #3357 from akirilov-arm/aarch64_emission_checks Cranelift AArch64: Avoid invalid encodings for some vector instructions	2021-09-17 12:48:18 -07:00
Nick Fitzgerald	a1f4b46f64	Bump Wasmtime to version 0.30.0; cranelift to 0.77.0	2021-09-17 10:33:50 -07:00
dheaton-arm	2f0ce4c86c	Implement `Smulhi` for interpreter Implemented `Smulhi` for the Cranelift interpreter, performing signed integer multiplication and producing the high half of a double-length result. Copyright (c) 2021, Arm Limited	2021-09-17 16:49:38 +01:00
dheaton-arm	83c3bc5b9d	Implement `Unarrow`, `Uunarrow`, and `Snarrow` for the interpreter Implemented the following Opcodes for the Cranelift interpreter: - `Unarrow` to combine two SIMD vectors into a new vector with twice the lanes but half the width, with signed inputs which are clamped to `0x00`. - `Uunarrow` to perform the same operation as `Unarrow` but treating inputs as unsigned. - `Snarrow` to perform the same operation as `Unarrow` but treating both inputs and outputs as signed, and saturating accordingly. Note that all 3 instructions saturate at the type boundaries. Copyright (c) 2021, Arm Limited	2021-09-17 13:26:10 +01:00
Anton Kirilov	a8aec2e0e6	Cranelift AArch64: Avoid invalid encodings for some vector instructions Copyright (c) 2021, Arm Limited.	2021-09-16 12:26:58 +01:00
Sam Parker	7da76f0601	cargo fmt	2021-09-15 16:01:51 +01:00
Sam Parker	80d596b055	AArch64 LSE atomic_rmw support Rename the existing AtomicRMW to AtomicRMWLoop and directly lower atomic_rmw operations, without a loop if LSE support is available. Copyright (c) 2021, Arm Limited	2021-09-15 16:01:51 +01:00
Chris Fallin	192586506d	Merge pull request #3342 from akirilov-arm/aarch64_lowering_type_checks Cranelift AArch64: Improve the type checks for IR operations	2021-09-13 10:12:06 -07:00
Anton Kirilov	8805e25042	Cranelift AArch64: Improve the type checks for IR operations There were cases where the AArch64 backend assumed that an IR operation would always operate on certain types (the most likely reason being that the corresponding WebAssembly instruction did not cover anything else), even though the definition of the IR operation imposed no constraints like that. Copyright (c) 2021, Arm Limited.	2021-09-13 14:46:45 +01:00
Afonso Bordado	92690b84a0	cranelift: Add SIMD `icmp` comparisons to interpreter	2021-09-11 17:15:44 +01:00
Ulrich Weigand	1b8154e0a3	cranelift: Fix big-endian regression in data_value.rs PR https://github.com/bytecodealliance/wasmtime/pull/3187 introduced a change to the write_to_slice and read_from_slice routines in data_value.rs that switched byte order on big-endian systems: the code used to use native byte order, and now hard-codes little-endian byte order. Fix by using native byte order again.	2021-09-11 15:06:25 +02:00
Afonso Bordado	3c1133379c	cranelift: Add `is_bool_vector` helper	2021-09-10 15:46:14 +01:00
Afonso Bordado	85d468dc5a	cranelift: Add `coerce_bools_to_ints` helper	2021-09-10 15:38:30 +01:00
Afonso Bordado	9460a4fb16	cranelift: Support bool vectors in trampoline	2021-09-10 15:10:51 +01:00
Chris Fallin	ecd795f736	Merge pull request #3290 from dheaton-arm/implement-ssatarith Implement `SaddSat` and `SsubSat` for the Cranelift interpreter	2021-09-03 09:48:34 -07:00
Chris Fallin	e3ccff0249	Merge pull request #3283 from dheaton-arm/implement-umulhi Implement `Umulhi` for the interpreter	2021-09-03 09:29:21 -07:00
dheaton-arm	8f057e0482	Implement `SaddSat` and `SsubSat` for the interpreter Implemented `SaddSat` and `SsubSat` to add and subtract signed vector values, saturating at the type boundaries rather than overflowing. Changed the parser to allow signed `i8` immediates in vectors as part of this work; fixes #3276. Copyright (c) 2021, Arm Limited.	2021-09-03 11:35:39 +01:00
dheaton-arm	562947c678	Fix CI tests + rename tests - Fixed CI tests for AArch64 and old x86. - Rename `simd-umulhi.clif` to `umulhi.clif`. - Rename `simd-umulhi-aarch64.clif` to `simd-umulhi.clif`. Copyright (c) 2021, Arm Limited.	2021-09-03 10:37:24 +01:00
Chris Fallin	6e05b646a3	Merge pull request #3282 from afonso360/x64-fix-brtables cranelift: Fix `br_table` for `i64` types in x64 backend.	2021-09-02 09:58:42 -07:00
Chris Fallin	2389a4ea00	Merge pull request #3274 from bnjbvr/fix-m1 A round of Mac M1 fixes	2021-09-02 09:49:01 -07:00
Chris Fallin	000a97f4ff	Merge pull request #3279 from dheaton-arm/implement-insertlane Implement `Insertlane` for the Cranelift interpreter	2021-09-02 09:44:59 -07:00
Afonso Bordado	f9ada24bcf	cranelift: Fix br_table for i64 inputs We still only support a maximum of u32::MAX entries, however we no longer crash when compiling 64 bit indexes. Fixes #3100	2021-09-02 15:31:48 +01:00
Benjamin Bouvier	85ec11acb9	Aarch64: always generate the CFA directive indicating no pointer signing	2021-09-02 09:16:34 +02:00
Benjamin Bouvier	fb94b81538	Use 16K code pages on Mac M1 Fixes #3278.	2021-09-02 09:16:34 +02:00
Alex Crichton	1532516a36	Use relative `call` instructions between wasm functions (#3275 ) * Use relative `call` instructions between wasm functions This commit is a relatively major change to the way that Wasmtime generates code for Wasm modules and how functions call each other. Prior to this commit all function calls between functions, even if they were defined in the same module, were done indirectly through a register. To implement this the backend would emit an absolute 8-byte relocation near all function calls, load that address into a register, and then call it. While this technique is simple to implement and easy to get right, it has two primary downsides associated with it: * Function calls are always indirect which means they are more difficult to predict, resulting in worse performance. * Generating a relocation-per-function call requires expensive relocation resolution at module-load time, which can be a large contributing factor to how long it takes to load a precompiled module. To fix these issues, while also somewhat compromising on the previously simple implementation technique, this commit switches wasm calls within a module to using the `colocated` flag enabled in Cranelift-speak, which basically means that a relative call instruction is used with a relocation that's resolved relative to the pc of the call instruction itself. When switching the `colocated` flag to `true` this commit is also then able to move much of the relocation resolution from `wasmtime_jit::link` into `wasmtime_cranelift::obj` during object-construction time. This frontloads all relocation work which means that there's actually no relocations related to function calls in the final image, solving both of our points above. The main gotcha in implementing this technique is that there are hardware limitations to relative function calls which mean we can't simply blindly use them. AArch64, for example, can only go +/- 64 MB from the `bl` instruction to the target, which means that if the function we're calling is a greater distance away then we would fail to resolve that relocation. On x86_64 the limits are +/- 2GB which are much larger, but theoretically still feasible to hit. Consequently the main increase in implementation complexity is fixing this issue. This issue is actually already present in Cranelift itself, and is internally one of the invariants handled by the `MachBuffer` type. When generating a function relative jumps between basic blocks have similar restrictions. This commit adds new methods for the `MachBackend` trait and updates the implementation of `MachBuffer` to account for all these new branches. Specifically the changes to `MachBuffer` are: * For AAarch64 the `LabelUse::Branch26` value now supports veneers, and AArch64 calls use this to resolve relocations. * The `emit_island` function has been rewritten internally to handle some cases which previously didn't come up before, such as: * When emitting an island the deadline is now recalculated, where previously it was always set to infinitely in the future. This was ok prior since only a `Branch19` supported veneers and once it was promoted no veneers were supported, so without multiple layers of promotion the lack of a new deadline was ok. * When emitting an island all pending fixups had veneers forced if their branch target wasn't known yet. This was generally ok for 19-bit fixups since the only kind getting a veneer was a 19-bit fixup, but with mixed kinds it's a bit odd to force veneers for a 26-bit fixup just because a nearby 19-bit fixup needed a veneer. Instead fixups are now re-enqueued unless they're known to be out-of-bounds. This may run the risk of generating more islands for 19-bit branches but it should also reduce the number of islands for between-function calls. * Otherwise the internal logic was tweaked to ideally be a bit more simple, but that's a pretty subjective criteria in compilers... I've added some simple testing of this for now. A synthetic compiler option was create to simply add padded 0s between functions and test cases implement various forms of calls that at least need veneers. A test is also included for x86_64, but it is unfortunately pretty slow because it requires generating 2GB of output. I'm hoping for now it's not too bad, but we can disable the test if it's prohibitive and otherwise just comment the necessary portions to be sure to run the ignored test if these parts of the code have changed. The final end-result of this commit is that for a large module I'm working with the number of relocations dropped to zero, meaning that nothing actually needs to be done to the text section when it's loaded into memory (yay!). I haven't run final benchmarks yet but this is the last remaining source of significant slowdown when loading modules, after I land a number of other PRs both active and ones that I only have locally for now. * Fix arm32 * Review comments	2021-09-01 13:27:38 -05:00
dheaton-arm	d956d349d8	Implement `Insertlane` for the Cranelift interpreter Implemented `Insertlane` to insert a value in the lane specified by the immediate value, overwriting the existing value in that lane. Added `TernaryImm8` support for the `imm_value` function. Copyright (c) 2021, Arm Limited.	2021-09-01 16:21:27 +01:00
bjorn3	a6598c310a	Remove empty preopt.serialized file It was added in #2312	2021-08-27 17:00:23 +02:00
bjorn3	690ea640b3	Implement fmin_pseudo and fmax_pseudo for scalars	2021-08-27 16:59:47 +02:00
Anton Kirilov	7b98be1bee	Cranelift: Simplify leaf functions that do not use the stack (#2960 ) * Cranelift AArch64: Simplify leaf functions that do not use the stack Leaf functions that do not use the stack (e.g. do not clobber any callee-saved registers) do not need a frame record. Copyright (c) 2021, Arm Limited.	2021-08-27 12:12:37 +02:00

1 2 3 4 5 ...

1426 Commits