Add support for BFloat16 - #821
Merged
Merged
Conversation
The JLL replaces its llvm-as product with llvm-downgrade, which reads bitcode instead of textual IR, and adds an llvm-dis-14 product. Bump the compat bound to 0.8 and feed metallib-as the parsed module as bitcode, targeting --bitcode-version=14.0. This matches the switch in GPUCompiler's mcgen that lets BFloat16 kernels compile on Julia 1.13+ (#817, #298). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Cover construction/conversion, arithmetic, abs/sqrt/min/max, and
reductions on MtlArray{BFloat16}, run on the device. Gated on
BFloat16s.llvm_arithmetic so they run only where Julia emits the native
`bfloat` IR type (Apple silicon needs LLVM 19+, i.e. Julia 1.13+);
elsewhere BFloat16 is an emulated i16, which we don't target here. These
complement GPUCompiler's AIR-lowering codegen tests with on-device
numerical checks (#298).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #821 +/- ##
==========================================
+ Coverage 81.34% 81.36% +0.02%
==========================================
Files 66 66
Lines 3318 3317 -1
==========================================
Hits 2699 2699
+ Misses 619 618 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Contributor
There was a problem hiding this comment.
Metal Benchmarks
Details
| Benchmark suite | Current: 0a8b1cd | Previous: bfb9ba3 | Ratio |
|---|---|---|---|
array/accumulate/Float32/1d |
813395.5 ns |
813334 ns |
1.00 |
array/accumulate/Float32/dims=1 |
999937.5 ns |
982750 ns |
1.02 |
array/accumulate/Float32/dims=1L |
10016375 ns |
10003208 ns |
1.00 |
array/accumulate/Float32/dims=2 |
1290479.5 ns |
1259354 ns |
1.02 |
array/accumulate/Float32/dims=2L |
6851375 ns |
5629958 ns |
1.22 |
array/accumulate/Int64/1d |
974208 ns |
950125 ns |
1.03 |
array/accumulate/Int64/dims=1 |
1168333 ns |
1125625.5 ns |
1.04 |
array/accumulate/Int64/dims=1L |
11953042 ns |
12141625 ns |
0.98 |
array/accumulate/Int64/dims=2 |
1449125 ns |
1476416 ns |
0.98 |
array/accumulate/Int64/dims=2L |
9439896 ns |
9438083 ns |
1.00 |
array/broadcast |
373042 ns |
374625 ns |
1.00 |
array/construct |
5834 ns |
5666 ns |
1.03 |
array/permutedims/2d |
643500 ns |
630750 ns |
1.02 |
array/permutedims/3d |
1127292 ns |
1117000 ns |
1.01 |
array/permutedims/4d |
1972625 ns |
1994209 ns |
0.99 |
array/private/copy |
427083.5 ns |
412292 ns |
1.04 |
array/private/copyto!/cpu_to_gpu |
365958 ns |
368583 ns |
0.99 |
array/private/copyto!/gpu_to_cpu |
354667 ns |
358916 ns |
0.99 |
array/private/copyto!/gpu_to_gpu |
338083 ns |
342666 ns |
0.99 |
array/private/iteration/findall/bool |
1080750 ns |
1073250 ns |
1.01 |
array/private/iteration/findall/int |
1255166 ns |
1252000 ns |
1.00 |
array/private/iteration/findfirst/bool |
1466041.5 ns |
1458437 ns |
1.01 |
array/private/iteration/findfirst/int |
1566750 ns |
1487958 ns |
1.05 |
array/private/iteration/findmin/1d |
1581916.5 ns |
1592041 ns |
0.99 |
array/private/iteration/findmin/2d |
1315667 ns |
1315875 ns |
1.00 |
array/private/iteration/logical |
1759125 ns |
1743542 ns |
1.01 |
array/private/iteration/scalar |
2519583.5 ns |
2638375.5 ns |
0.95 |
array/random/rand/Float32 |
601271 ns |
634917 ns |
0.95 |
array/random/rand/Int64 |
698167 ns |
669834 ns |
1.04 |
array/random/rand!/Float32 |
572167 ns |
580958 ns |
0.98 |
array/random/rand!/Int64 |
504334 ns |
509000 ns |
0.99 |
array/random/randn/Float32 |
595834 ns |
597958 ns |
1.00 |
array/random/randn!/Float32 |
519375 ns |
531209 ns |
0.98 |
array/reductions/mapreduce/Float32/1d |
505667 ns |
750833 ns |
0.67 |
array/reductions/mapreduce/Float32/dims=1 |
509792 ns |
499041.5 ns |
1.02 |
array/reductions/mapreduce/Float32/dims=1L |
828291.5 ns |
780791 ns |
1.06 |
array/reductions/mapreduce/Float32/dims=2 |
504458 ns |
502750 ns |
1.00 |
array/reductions/mapreduce/Float32/dims=2L |
1352000 ns |
1356041 ns |
1.00 |
array/reductions/mapreduce/Int64/1d |
959041 ns |
934917 ns |
1.03 |
array/reductions/mapreduce/Int64/dims=1 |
806709 ns |
786666 ns |
1.03 |
array/reductions/mapreduce/Int64/dims=1L |
1547125 ns |
1712500 ns |
0.90 |
array/reductions/mapreduce/Int64/dims=2 |
996292 ns |
966667 ns |
1.03 |
array/reductions/mapreduce/Int64/dims=2L |
2235750 ns |
2260917 ns |
0.99 |
array/reductions/reduce/Float32/1d |
499583 ns |
743333 ns |
0.67 |
array/reductions/reduce/Float32/dims=1 |
503500 ns |
499625 ns |
1.01 |
array/reductions/reduce/Float32/dims=1L |
828521 ns |
813625 ns |
1.02 |
array/reductions/reduce/Float32/dims=2 |
504417 ns |
505833 ns |
1.00 |
array/reductions/reduce/Float32/dims=2L |
1349750 ns |
1346625 ns |
1.00 |
array/reductions/reduce/Int64/1d |
949708 ns |
930625 ns |
1.02 |
array/reductions/reduce/Int64/dims=1 |
786417 ns |
783875 ns |
1.00 |
array/reductions/reduce/Int64/dims=1L |
1708520.5 ns |
1680125 ns |
1.02 |
array/reductions/reduce/Int64/dims=2 |
997125 ns |
980563 ns |
1.02 |
array/reductions/reduce/Int64/dims=2L |
2233000 ns |
2260250 ns |
0.99 |
array/shared/copy |
218583 ns |
238375 ns |
0.92 |
array/shared/copyto!/cpu_to_gpu |
40917 ns |
40667 ns |
1.01 |
array/shared/copyto!/gpu_to_cpu |
39917 ns |
40667 ns |
0.98 |
array/shared/copyto!/gpu_to_gpu |
41041 ns |
41292 ns |
0.99 |
array/shared/iteration/findall/bool |
1090375 ns |
1079166 ns |
1.01 |
array/shared/iteration/findall/int |
1255209 ns |
1250333 ns |
1.00 |
array/shared/iteration/findfirst/bool |
1195062.5 ns |
1192416.5 ns |
1.00 |
array/shared/iteration/findfirst/int |
1232208 ns |
1274104.5 ns |
0.97 |
array/shared/iteration/findmin/1d |
1326708 ns |
1282291 ns |
1.03 |
array/shared/iteration/findmin/2d |
1319042 ns |
1266125 ns |
1.04 |
array/shared/iteration/logical |
1607417 ns |
1594459 ns |
1.01 |
array/shared/iteration/scalar |
5847.166666666667 ns |
5868.166666666667 ns |
1.00 |
integration/byval/reference |
1159291 ns |
1157708 ns |
1.00 |
integration/byval/slices=1 |
1160625 ns |
1159208 ns |
1.00 |
integration/byval/slices=2 |
2088958 ns |
2086791.5 ns |
1.00 |
integration/byval/slices=3 |
9209792 ns |
7931979 ns |
1.16 |
integration/metaldevrt |
465875 ns |
468520.5 ns |
0.99 |
kernel/indexing |
360291 ns |
366500 ns |
0.98 |
kernel/indexing_checked |
539875 ns |
540834 ns |
1.00 |
kernel/launch |
13500 ns |
13208 ns |
1.02 |
kernel/rand |
545500 ns |
558042 ns |
0.98 |
latency/import |
1407926812 ns |
1399005208 ns |
1.01 |
latency/precompile |
31342379104 ns |
31215035583.5 ns |
1.00 |
latency/ttfp |
1709751979.5 ns |
1710418937.5 ns |
1.00 |
metal/synchronization/context |
814.3203883495146 ns |
838.258064516129 ns |
0.97 |
metal/synchronization/stream |
431.5326633165829 ns |
436.9748743718593 ns |
0.99 |
This comment was automatically generated by workflow using github-action-benchmark.
BFloat16 works on Metal across every supported Julia: an i16 emulation pre-1.13 (arithmetic widened to Float32 by BFloat16s) and the native `bfloat` type on 1.13+. Run the execution tests on both paths rather than gating them to the native one, since the emulated path is correct too. Add codegen tests, gated to the native path, that assert a real BFloat16 kernel compiles to `bfloat` IR (`fadd bfloat`) and that abs reaches AIR as the f32-promoted builtin (`air.fabs.f32`) it lacks for bfloat. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Member
Author
|
Turns out that software-emulated BFloat16 already worked. I wonder if we should warn users about this, but it's tricky, since our current bad datatype detection operates at the LLVM level (and software-emulated BFLoat16 is simply |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Julia 1.13+ only, since that's the only release where we generate
bfloaton AArch64 (even though this is only a property of the CPU back-end...).Fixes #298
Fixes #817