wasmtime

Author	SHA1	Message	Date
Afonso Bordado	20913a7473	cranelift: Enable compiling br_tables for types larger than i32	2021-09-02 16:26:23 +01:00
Alex Crichton	37b9fc5333	Fix async build	2021-09-02 07:38:29 -07:00
Afonso Bordado	f9ada24bcf	cranelift: Fix br_table for i64 inputs We still only support a maximum of u32::MAX entries, however we no longer crash when compiling 64 bit indexes. Fixes #3100	2021-09-02 15:31:48 +01:00
Alex Crichton	6b5e21d80e	Inline some trivial store accessors These were showing up in some profiles, but they're trivial functions, so `#[inline]` them.	2021-09-02 07:26:10 -07:00
Alex Crichton	230159efa7	Inline some type conversions for `()` The `()` type accidentally wasn't getting its trivial type conversions inlined because it doesn't actually have any type parameters. This commit adds `#[inline]` to the relevant functions to ensure that these get inlined across crates.	2021-09-02 07:26:07 -07:00
Alex Crichton	c8f55ed688	Optimize codegen slightly calling wasm functions Currently wasm-calls work with `Result<T, Trap>` internally but `Trap` is an enum defined in `wasmtime-runtime` which is actually quite large. Since traps are supposed to be rare this commit changes these functions to return a `Box<Trap>` which is un-boxed later up in the `wasmtime` crate within a `#[cold]` function.	2021-09-02 07:26:03 -07:00
Benjamin Bouvier	0cf9b2d5e6	Remove cargo check for old rustc versions Firefox has updated its toolchain to something much more recent.	2021-09-02 16:06:02 +02:00
Alex Crichton	25755ff23a	Discuss precompiled modules at today's meeting (#3280 ) A new agenda item!	2021-09-02 09:03:54 -05:00
dheaton-arm	16b6a404e4	Implement `Umulhi` for the interpreter Implemented `Umulhi` for the Cranelift interpreter, performing unsigned integer multiplication and producing the high half of a double-length result. Fixed `ExtractUpper` conversion behaviour as part of this change, which was extracting from a 128-bit value regardless of the size of the original value. Copyright (c) 2021, Arm Limited.	2021-09-02 13:11:41 +01:00
Benjamin Bouvier	85ec11acb9	Aarch64: always generate the CFA directive indicating no pointer signing	2021-09-02 09:16:34 +02:00
Benjamin Bouvier	fb94b81538	Use 16K code pages on Mac M1 Fixes #3278.	2021-09-02 09:16:34 +02:00
Benjamin Bouvier	f871e8cf8f	Correctly set the address of FP when unwinding from within fibers on aarch64 Fixes #3256.	2021-09-02 08:58:03 +02:00
Pat Hickey	f46f58ecc2	replace Config::deserialize_check_wasmtime_version with Config::module_version which is more expressive than the former. Instead of just configuring Module::deserialize to ignore version information, we can configure Module::serialize to emit a custom version string, and Module::deserialize to check for that string. A new enum ModuleVersionStrategy is declared, and Config::deserialize_check_wasmtime_version:bool is replaced with Config::module_version:ModuleVersionStrategy.	2021-09-01 17:12:15 -07:00
Afonso Bordado	f0f2efba26	cranelift: CLIF Fuzzer generate brz/brnz/bricmp instructions	2021-09-01 19:44:11 +01:00
Afonso Bordado	f4bd7d17a3	cranelift: CLIF Fuzzer generate multiple blocks	2021-09-01 19:44:05 +01:00
Afonso Bordado	fb1201cb60	cranelift: CLIF Fuzzer limit instructions executed in interpreter	2021-09-01 19:43:58 +01:00
Alex Crichton	1532516a36	Use relative `call` instructions between wasm functions (#3275 ) * Use relative `call` instructions between wasm functions This commit is a relatively major change to the way that Wasmtime generates code for Wasm modules and how functions call each other. Prior to this commit all function calls between functions, even if they were defined in the same module, were done indirectly through a register. To implement this the backend would emit an absolute 8-byte relocation near all function calls, load that address into a register, and then call it. While this technique is simple to implement and easy to get right, it has two primary downsides associated with it: * Function calls are always indirect which means they are more difficult to predict, resulting in worse performance. * Generating a relocation-per-function call requires expensive relocation resolution at module-load time, which can be a large contributing factor to how long it takes to load a precompiled module. To fix these issues, while also somewhat compromising on the previously simple implementation technique, this commit switches wasm calls within a module to using the `colocated` flag enabled in Cranelift-speak, which basically means that a relative call instruction is used with a relocation that's resolved relative to the pc of the call instruction itself. When switching the `colocated` flag to `true` this commit is also then able to move much of the relocation resolution from `wasmtime_jit::link` into `wasmtime_cranelift::obj` during object-construction time. This frontloads all relocation work which means that there's actually no relocations related to function calls in the final image, solving both of our points above. The main gotcha in implementing this technique is that there are hardware limitations to relative function calls which mean we can't simply blindly use them. AArch64, for example, can only go +/- 64 MB from the `bl` instruction to the target, which means that if the function we're calling is a greater distance away then we would fail to resolve that relocation. On x86_64 the limits are +/- 2GB which are much larger, but theoretically still feasible to hit. Consequently the main increase in implementation complexity is fixing this issue. This issue is actually already present in Cranelift itself, and is internally one of the invariants handled by the `MachBuffer` type. When generating a function relative jumps between basic blocks have similar restrictions. This commit adds new methods for the `MachBackend` trait and updates the implementation of `MachBuffer` to account for all these new branches. Specifically the changes to `MachBuffer` are: * For AAarch64 the `LabelUse::Branch26` value now supports veneers, and AArch64 calls use this to resolve relocations. * The `emit_island` function has been rewritten internally to handle some cases which previously didn't come up before, such as: * When emitting an island the deadline is now recalculated, where previously it was always set to infinitely in the future. This was ok prior since only a `Branch19` supported veneers and once it was promoted no veneers were supported, so without multiple layers of promotion the lack of a new deadline was ok. * When emitting an island all pending fixups had veneers forced if their branch target wasn't known yet. This was generally ok for 19-bit fixups since the only kind getting a veneer was a 19-bit fixup, but with mixed kinds it's a bit odd to force veneers for a 26-bit fixup just because a nearby 19-bit fixup needed a veneer. Instead fixups are now re-enqueued unless they're known to be out-of-bounds. This may run the risk of generating more islands for 19-bit branches but it should also reduce the number of islands for between-function calls. * Otherwise the internal logic was tweaked to ideally be a bit more simple, but that's a pretty subjective criteria in compilers... I've added some simple testing of this for now. A synthetic compiler option was create to simply add padded 0s between functions and test cases implement various forms of calls that at least need veneers. A test is also included for x86_64, but it is unfortunately pretty slow because it requires generating 2GB of output. I'm hoping for now it's not too bad, but we can disable the test if it's prohibitive and otherwise just comment the necessary portions to be sure to run the ignored test if these parts of the code have changed. The final end-result of this commit is that for a large module I'm working with the number of relocations dropped to zero, meaning that nothing actually needs to be done to the text section when it's loaded into memory (yay!). I haven't run final benchmarks yet but this is the last remaining source of significant slowdown when loading modules, after I land a number of other PRs both active and ones that I only have locally for now. * Fix arm32 * Review comments	2021-09-01 13:27:38 -05:00
Chris Fallin	91410aaddf	Merge pull request #3234 from dheaton-arm/implement-isubb Implement `IsubBin`, `IsubBout`, and `IsubBorrow`for Cranelift interpreter	2021-09-01 11:25:43 -07:00
Chris Fallin	2a63979151	Merge pull request #3258 from afonso360/ssa-dce cranelift: Prevent infinite loops in ssa frontend with unreachable code.	2021-09-01 11:19:40 -07:00
Aaron Turner	191051b644	Docs: Created the Wasmtime Markdown Parser Example (#3193 ) * Finished the Markdown Parser Example for Wasmtime * Made requested changes * Tiny change to explanation of `--dir` CLI arg * Add `bash` annotations to shell script code blocks * Trying to fix the markdown example bug * Figured out rustdoc, and what needed to be done * Made requested changes Co-authored-by: Till Schneidereit <till@tillschneidereit.net>	2021-09-01 12:25:36 -05:00
dheaton-arm	d956d349d8	Implement `Insertlane` for the Cranelift interpreter Implemented `Insertlane` to insert a value in the lane specified by the immediate value, overwriting the existing value in that lane. Added `TernaryImm8` support for the `imm_value` function. Copyright (c) 2021, Arm Limited.	2021-09-01 16:21:27 +01:00
dheaton-arm	4cdb2d3dac	Merge vector iterators into chain Copyrght (c) 2021, Arm Limited	2021-09-01 15:55:35 +01:00
Afonso Bordado	f9f5ae59a6	cranelift: Merge interpreter tests with runtests (#3252 ) Almost all the tests in the interpreter are already in the runtests folder so that we can reuse them for the backends. The distinction between interpreter tests and runtests is no longer very clear, since they should both support the same clif code, and produce the same results. We only have two test files: * `add.clif` tests the add and jump instruction, both of which are already covered in other test files, so we remove that file. * `fibonacci.clif` does a recursive call which is currently not supported in the filetest environment, so we keep this test interpreter only for now.	2021-09-01 06:42:02 -07:00
dheaton-arm	7a5646c5f4	Implement `IaddPairwise` for the interpreter Implemented `IaddPairwise` for the Cranelift interpreter, to add pairs of adjacent values in two SIMD vectors, concatenating them at the end (preserving both lane size and number of lanes). Copyright (c) 2021, Arm Limited	2021-09-01 13:53:26 +01:00
Afonso Bordado	6a9378e244	cranelift: Prevent infinite loops in ssa frontend with unreachable code. Perform a search over block predecessors trying to find loops of unreachable predecessors. We do this by iterating on predecessors and marking them as visited, stopping if we find a previously visited block or if we find a block with multiple predecessors. This issue was found by the CLIF fuzzer in #3094.	2021-09-01 11:19:23 +01:00
Chris Fallin	5843d53481	Merge pull request #3270 from sunfishcode/sunfishcode/use-rust-alloc Use `std::alloc::alloc` instead of `libc::posix_memalign`.	2021-08-31 18:51:40 -07:00
Dan Gohman	05d113148d	Use `std::alloc::alloc` instead of `libc::posix_memalign`. This makes Cranelift use the Rust `alloc` API its allocations, rather than directly calling into `libc`, which makes it respect the `#[global_allocator]` configuration. Also, use `region::page::ceil` instead of having our own copies of that logic.	2021-08-31 15:49:50 -07:00
Dan Gohman	197aec9a08	Update io-lifetimes, cap-std, and rsix (#3269 ) - Fixes for compiling on OpenBSD - io-lifetimes 0.3.0 has an option (io_lifetimes_use_std, which is off by default) for testing the `io_safety` feature in Rust nightly.	2021-08-31 13:02:37 -07:00
Alex Crichton	9e0c910023	Add a `Module::deserialize_file` method (#3266 ) * Add a `Module::deserialize_file` method This commit adds a new method to the `wasmtime::Module` type, `deserialize_file`. This is intended to be the same as the `deserialize` method except for the serialized module is present as an on-disk file. This enables Wasmtime to internally use `mmap` to avoid copying bytes around and generally makes loading a module much faster. A C API is added in this commit as well for various bindings to use this accelerated path now as well. Another option perhaps for a Rust-based API is to have an API taking a `File` itself to allow for a custom file descriptor in one way or another, but for now that's left for a possible future refactoring if we find a use case. * Fix compat with main - handle readdonly mmap * wip * Try to fix Windows support	2021-08-31 13:05:51 -05:00
Damian Heaton	4378ea8e01	Implement `IaddCin`, `IaddCout`, and `IaddCarry` for Cranelift interpreter (#3233 ) * Implement `IaddCin`, `IaddCout`, and `IaddCarry` for Cranelift interpreter Implemented the following Opcodes for the Cranelift interpreter: - `IaddCin` to add two scalar integers with an input carry flag. - `IaddCout` to add two scalar integers and report overflow with the carry flag. - `IaddCarry` to add two scalar integers with an input carry flag, reporting overflow with the output carry flag. Copyright (c) 2021, Arm Limited * Simplify carry check + add i64 `IaddCarry` tests Copyright (c) 2021, Arm Limited * Move tests to `runtests` Copyright (c) 2021, Arm Limited	2021-08-31 09:29:38 -07:00
Alex Crichton	4376cf2609	Add differential fuzzing against V8 (#3264 ) * Add differential fuzzing against V8 This commit adds a differential fuzzing target to Wasmtime along the lines of the wasmi and spec interpreters we already have, but with V8 instead. The intention here is that wasmi is unlikely to receive updates over time (e.g. for SIMD), and the spec interpreter is not suitable for fuzzing against in general due to its performance characteristics. The hope is that V8 is indeed appropriate to fuzz against because it's naturally receiving updates and it also is expected to have good performance. Here the `rusty_v8` crate is used which provides bindings to V8 as well as precompiled binaries by default. This matches exactly the use case we need and at least for now I think the `rusty_v8` crate will be maintained by the Deno folks as they continue to develop it. If it becomes an issue though maintaining we can evaluate other options to have differential fuzzing against. For now this commit enables the SIMD and bulk-memory feature of fuzz-target-generation which should enable them to get differentially-fuzzed with V8 in addition to the compilation fuzzing we're already getting. * Use weak linkage for GDB jit helpers This should help us deduplicate our symbol with other JIT runtimes, if any. For now this leans on some C helpers to define the weak linkage since Rust doesn't support that on stable yet. * Don't use rusty_v8 on MinGW They don't have precompiled libraries there. * Fix msvc build * Comment about execution	2021-08-31 09:34:55 -05:00
dheaton-arm	d1fe72affa	Add `i64` tests to `IsubBorrow` and move tests. Copyright (c) 2021, Arm Limited	2021-08-31 11:47:26 +01:00
Alex Crichton	ef3ec594ce	Don't copy executable code into a `CodeMemory` (#3265 ) * Don't copy executable code into a `CodeMemory` This commit moves a copy from compiled artifacts into a `CodeMemory`. In general this commit drastically changes the meaning of a `CodeMemory`. Previously it was an iteratively-pushed-on structure that would accumulate executable code over time. Afterwards, however, it's a manager for an `MmapVec` which updates the permissions on text section to ensure that the pages are executable. By taking ownership of an `MmapVec` within a `CodeMemory` there's no need to copy any data around, which means that the `.text` section in the ELF image produced by Wasmtime is usable as-is after placement in memory and relocations have been resolved. This moves Wasmtime one step closer to being able to directly use a module after it's `mmap`'d into memory, optimizing when a module is loaded. * Fix windows section alignment * Review comments	2021-08-30 13:38:35 -05:00
Alex Crichton	eb251deca9	Remove `scroll` dependency from `wasmtime-jit` (#3260 ) Similar functionality to `scroll` is provided with the `object` crate and doesn't have a `*_derive` crate to go with it. This commit updates the jitdump linux support to use `object` instead of `scroll` to achieve the needs of writing structs-as-bytes onto disk.	2021-08-30 13:26:07 -05:00
Alex Crichton	a978c7e384	Update wasm-smith (#3267 ) Brings in a fix for a fuzz-bug found on oss-fuzz.	2021-08-30 11:48:50 -05:00
Nick Fitzgerald	1c8f0b4652	Merge pull request #3261 from jlb6740/fix-build-for-benchmark-api Bench-api cargo update to allow seeing Module functions	2021-08-30 09:39:49 -07:00
Alex Crichton	a237e73b5a	Remove some allocations in `CodeMemory` (#3253 ) * Remove some allocations in `CodeMemory` This commit removes the `FinishedFunctions` type as well as allocations associated with trampolines when allocating inside of a `CodeMemory`. The main goal of this commit is to improve the time spent in `CodeMemory` where currently today a good portion of time is spent simply parsing symbol names and trying to extract function indices from them. Instead this commit implements a new strategy (different from #3236) where compilation records offset/length information for all functions/trampolines so this doesn't need to be re-learned from the object file later. A consequence of this commit is that this offset information will be decoded/encoded through `bincode` unconditionally, but we can also optimize that later if necessary as well. Internally this involved quite a bit of refactoring since the previous map for `FinishedFunctions` was relatively heavily relied upon. * comments	2021-08-30 10:35:17 -05:00
Alex Crichton	c73be1f13a	Use an mmap-friendly serialization format (#3257 ) * Use an mmap-friendly serialization format This commit reimplements the main serialization format for Wasmtime's precompiled artifacts. Previously they were generally a binary blob of `bincode`-encoded metadata prefixed with some versioning information. The downside of this format, though, is that loading a precompiled artifact required pushing all information through `bincode`. This is inefficient when some data, such as trap/address tables, are rarely accessed. The new format added in this commit is one which is designed to be `mmap`-friendly. This means that the relevant parts of the precompiled artifact are already page-aligned for updating permissions of pieces here and there. Additionally the artifact is optimized so that if data is rarely read then we can delay reading it until necessary. The new artifact format for serialized modules is an ELF file. This is not a public API guarantee, so it cannot be relied upon. In the meantime though this is quite useful for exploring precompiled modules with standard tooling like `objdump`. The ELF file is already constructed as part of module compilation, and this is the main contents of the serialized artifact. THere is some extra information, though, not encoded in each module's individual ELF file such as type information. This information continues to be `bincode`-encoded, but it's intended to be much smaller and much faster to deserialize. This extra information is appended to the end of the ELF file. This means that the original ELF file is still a valid ELF file, we just get to have extra bits at the end. More information on the new format can be found in the module docs of the serialization module of Wasmtime. Another refatoring implemented as part of this commit is to deserialize and store object files directly in `mmap`-backed storage. This avoids the need to copy bytes after the artifact is loaded into memory for each compiled module, and in a future commit it opens up the door to avoiding copying the text section into a `CodeMemory`. For now, though, the main change is that copies are not necessary when loading from a precompiled compilation artifact once the artifact is itself in mmap-based memory. To assist with managing `mmap`-based memory a new `MmapVec` type was added to `wasmtime_jit` which acts as a form of `Vec<T>` backed by a `wasmtime_runtime::Mmap`. This type notably supports `drain(..N)` to slice the buffer into disjoint regions that are all separately owned, such as having a separately owned window into one artifact for all object files contained within. Finally this commit implements a small refactoring in `wasmtime-cache` to use the standard artifact format for cache entries rather than a bincode-encoded version. This required some more hooks for serializing/deserializing but otherwise the crate still performs as before. * Review comments	2021-08-30 09:19:20 -05:00
Chris Fallin	16854e73c5	Merge pull request #3115 from bjorn3/fminmax_pseudo_scalar Implement fmin_pseudo and fmax_pseudo for scalars	2021-08-29 19:07:01 -07:00
Johnnie Birch	6e1015c0b6	Bench-api cargo update to allow seeing Module functions	2021-08-28 12:41:13 -07:00
Andrew Brown	4ccdcb110a	typo: change 'sharedable' to 'shareable' (#3259 )	2021-08-27 11:50:11 -07:00
bjorn3	b79e59882d	Fix tests	2021-08-27 18:28:33 +02:00
bjorn3	8adb40b2b8	Add tests	2021-08-27 17:48:04 +02:00
bjorn3	a6598c310a	Remove empty preopt.serialized file It was added in #2312	2021-08-27 17:00:23 +02:00
bjorn3	690ea640b3	Implement fmin_pseudo and fmax_pseudo for scalars	2021-08-27 16:59:47 +02:00
Anton Kirilov	7b98be1bee	Cranelift: Simplify leaf functions that do not use the stack (#2960 ) * Cranelift AArch64: Simplify leaf functions that do not use the stack Leaf functions that do not use the stack (e.g. do not clobber any callee-saved registers) do not need a frame record. Copyright (c) 2021, Arm Limited.	2021-08-27 12:12:37 +02:00
Alex Crichton	12515e6646	Move trap information to a section of the compiled image (#3241 ) This commit moves the `traps` field of `FunctionInfo` into a section of the compiled artifact produced by Cranelift. This section is quite large and when previously encoded/decoded with `bincode` this can take quite some time to process. Traps are expected to be relatively rare and it's not necessarily the right tradeoff to spend so much time serializing/deserializing this data, so this commit offloads the section into a custom-encoded binary format located elsewhere in the compiled image. This is similar to #3240 in its goal which is to move very large pieces of metadata to their own sections to avoid decoding anything when we load a precompiled modules. This also has a small benefit that it's slightly more efficient storage for the trap information too, but that's a negligible benefit. This is part of #3230 to make loading modules fast.	2021-08-27 01:09:55 -05:00
Alex Crichton	fc91176685	Move address maps to a section of the compiled image (#3240 ) This commit moves the `address_map` field of `FunctionInfo` into a custom-encoded section of the executable. The goal of this commit is, as previous commits, to push less data through `bincode`. The `address_map` field is actually extremely large and has huge benefits of not being decoded when we load a module. This data is only used for traps and such as well, so it's not overly important that it's massaged in to precise data the runtime can extremely speedily use. The `FunctionInfo` type does retain a tiny bit of information about the function itself (it's start source location), but other than that the `FunctionAddressMap` structure is moved from `wasmtime-environ` to `wasmtime-cranelift` since it's now no longer needed outside of that context.	2021-08-26 23:06:41 -05:00
Alex Crichton	d12f1d77e6	Convert compilation artifacts to just bytes (#3239 ) * Convert compilation artifacts to just bytes This commit strips the `CompilationArtifacts` type down to simply a list of bytes. This moves all extra metadata elsewhere to live within the list of bytes itself as `bincode`-encoded information. Small affordance is made to avoid an in-process serialize-then-deserialize round-trip for use cases like `Module::new`, but otherwise this is mostly just moving some data around. * Rename data section to `.rodata.wasm`	2021-08-26 21:17:02 -05:00
Peter Huene	a2a6be72c4	Merge pull request #3245 from peterhuene/add-paged-init-setting Add `paged_memory_initialization` to Config.	2021-08-26 18:54:16 -07:00

... 2 3 4 5 6 ...

8918 Commits