Skip to content

Add support for BFloat16 - #821

Merged
maleadt merged 4 commits into
mainfrom
tb/bfloat16
Jun 9, 2026
Merged

maleadt merged 4 commits into
mainfrom
tb/bfloat16

Conversation

@maleadt

@maleadt maleadt commented Jun 9, 2026

Copy link
Copy Markdown
Member

Julia 1.13+ only, since that's the only release where we generate bfloat on AArch64 (even though this is only a property of the CPU back-end...).

Fixes #298
Fixes #817

maleadt and others added 3 commits June 9, 2026 14:12
The JLL replaces its llvm-as product with llvm-downgrade, which reads
bitcode instead of textual IR, and adds an llvm-dis-14 product. Bump the
compat bound to 0.8 and feed metallib-as the parsed module as bitcode,
targeting --bitcode-version=14.0. This matches the switch in GPUCompiler's
mcgen that lets BFloat16 kernels compile on Julia 1.13+ (#817, #298).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Cover construction/conversion, arithmetic, abs/sqrt/min/max, and
reductions on MtlArray{BFloat16}, run on the device. Gated on
BFloat16s.llvm_arithmetic so they run only where Julia emits the native
`bfloat` IR type (Apple silicon needs LLVM 19+, i.e. Julia 1.13+);
elsewhere BFloat16 is an emulated i16, which we don't target here. These
complement GPUCompiler's AIR-lowering codegen tests with on-device
numerical checks (#298).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@codecov

codecov Bot commented Jun 9, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 81.36%. Comparing base (0d8f897) to head (0a8b1cd).
⚠️ Report is 8 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #821      +/-   ##
==========================================
+ Coverage   81.34%   81.36%   +0.02%     
==========================================
  Files          66       66              
  Lines        3318     3317       -1     
==========================================
  Hits         2699     2699              
+ Misses        619      618       -1     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions github-actions Bot left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Metal Benchmarks

Details
Benchmark suite Current: 0a8b1cd Previous: bfb9ba3 Ratio
array/accumulate/Float32/1d 813395.5 ns 813334 ns 1.00
array/accumulate/Float32/dims=1 999937.5 ns 982750 ns 1.02
array/accumulate/Float32/dims=1L 10016375 ns 10003208 ns 1.00
array/accumulate/Float32/dims=2 1290479.5 ns 1259354 ns 1.02
array/accumulate/Float32/dims=2L 6851375 ns 5629958 ns 1.22
array/accumulate/Int64/1d 974208 ns 950125 ns 1.03
array/accumulate/Int64/dims=1 1168333 ns 1125625.5 ns 1.04
array/accumulate/Int64/dims=1L 11953042 ns 12141625 ns 0.98
array/accumulate/Int64/dims=2 1449125 ns 1476416 ns 0.98
array/accumulate/Int64/dims=2L 9439896 ns 9438083 ns 1.00
array/broadcast 373042 ns 374625 ns 1.00
array/construct 5834 ns 5666 ns 1.03
array/permutedims/2d 643500 ns 630750 ns 1.02
array/permutedims/3d 1127292 ns 1117000 ns 1.01
array/permutedims/4d 1972625 ns 1994209 ns 0.99
array/private/copy 427083.5 ns 412292 ns 1.04
array/private/copyto!/cpu_to_gpu 365958 ns 368583 ns 0.99
array/private/copyto!/gpu_to_cpu 354667 ns 358916 ns 0.99
array/private/copyto!/gpu_to_gpu 338083 ns 342666 ns 0.99
array/private/iteration/findall/bool 1080750 ns 1073250 ns 1.01
array/private/iteration/findall/int 1255166 ns 1252000 ns 1.00
array/private/iteration/findfirst/bool 1466041.5 ns 1458437 ns 1.01
array/private/iteration/findfirst/int 1566750 ns 1487958 ns 1.05
array/private/iteration/findmin/1d 1581916.5 ns 1592041 ns 0.99
array/private/iteration/findmin/2d 1315667 ns 1315875 ns 1.00
array/private/iteration/logical 1759125 ns 1743542 ns 1.01
array/private/iteration/scalar 2519583.5 ns 2638375.5 ns 0.95
array/random/rand/Float32 601271 ns 634917 ns 0.95
array/random/rand/Int64 698167 ns 669834 ns 1.04
array/random/rand!/Float32 572167 ns 580958 ns 0.98
array/random/rand!/Int64 504334 ns 509000 ns 0.99
array/random/randn/Float32 595834 ns 597958 ns 1.00
array/random/randn!/Float32 519375 ns 531209 ns 0.98
array/reductions/mapreduce/Float32/1d 505667 ns 750833 ns 0.67
array/reductions/mapreduce/Float32/dims=1 509792 ns 499041.5 ns 1.02
array/reductions/mapreduce/Float32/dims=1L 828291.5 ns 780791 ns 1.06
array/reductions/mapreduce/Float32/dims=2 504458 ns 502750 ns 1.00
array/reductions/mapreduce/Float32/dims=2L 1352000 ns 1356041 ns 1.00
array/reductions/mapreduce/Int64/1d 959041 ns 934917 ns 1.03
array/reductions/mapreduce/Int64/dims=1 806709 ns 786666 ns 1.03
array/reductions/mapreduce/Int64/dims=1L 1547125 ns 1712500 ns 0.90
array/reductions/mapreduce/Int64/dims=2 996292 ns 966667 ns 1.03
array/reductions/mapreduce/Int64/dims=2L 2235750 ns 2260917 ns 0.99
array/reductions/reduce/Float32/1d 499583 ns 743333 ns 0.67
array/reductions/reduce/Float32/dims=1 503500 ns 499625 ns 1.01
array/reductions/reduce/Float32/dims=1L 828521 ns 813625 ns 1.02
array/reductions/reduce/Float32/dims=2 504417 ns 505833 ns 1.00
array/reductions/reduce/Float32/dims=2L 1349750 ns 1346625 ns 1.00
array/reductions/reduce/Int64/1d 949708 ns 930625 ns 1.02
array/reductions/reduce/Int64/dims=1 786417 ns 783875 ns 1.00
array/reductions/reduce/Int64/dims=1L 1708520.5 ns 1680125 ns 1.02
array/reductions/reduce/Int64/dims=2 997125 ns 980563 ns 1.02
array/reductions/reduce/Int64/dims=2L 2233000 ns 2260250 ns 0.99
array/shared/copy 218583 ns 238375 ns 0.92
array/shared/copyto!/cpu_to_gpu 40917 ns 40667 ns 1.01
array/shared/copyto!/gpu_to_cpu 39917 ns 40667 ns 0.98
array/shared/copyto!/gpu_to_gpu 41041 ns 41292 ns 0.99
array/shared/iteration/findall/bool 1090375 ns 1079166 ns 1.01
array/shared/iteration/findall/int 1255209 ns 1250333 ns 1.00
array/shared/iteration/findfirst/bool 1195062.5 ns 1192416.5 ns 1.00
array/shared/iteration/findfirst/int 1232208 ns 1274104.5 ns 0.97
array/shared/iteration/findmin/1d 1326708 ns 1282291 ns 1.03
array/shared/iteration/findmin/2d 1319042 ns 1266125 ns 1.04
array/shared/iteration/logical 1607417 ns 1594459 ns 1.01
array/shared/iteration/scalar 5847.166666666667 ns 5868.166666666667 ns 1.00
integration/byval/reference 1159291 ns 1157708 ns 1.00
integration/byval/slices=1 1160625 ns 1159208 ns 1.00
integration/byval/slices=2 2088958 ns 2086791.5 ns 1.00
integration/byval/slices=3 9209792 ns 7931979 ns 1.16
integration/metaldevrt 465875 ns 468520.5 ns 0.99
kernel/indexing 360291 ns 366500 ns 0.98
kernel/indexing_checked 539875 ns 540834 ns 1.00
kernel/launch 13500 ns 13208 ns 1.02
kernel/rand 545500 ns 558042 ns 0.98
latency/import 1407926812 ns 1399005208 ns 1.01
latency/precompile 31342379104 ns 31215035583.5 ns 1.00
latency/ttfp 1709751979.5 ns 1710418937.5 ns 1.00
metal/synchronization/context 814.3203883495146 ns 838.258064516129 ns 0.97
metal/synchronization/stream 431.5326633165829 ns 436.9748743718593 ns 0.99

This comment was automatically generated by workflow using github-action-benchmark.

BFloat16 works on Metal across every supported Julia: an i16 emulation
pre-1.13 (arithmetic widened to Float32 by BFloat16s) and the native
`bfloat` type on 1.13+. Run the execution tests on both paths rather than
gating them to the native one, since the emulated path is correct too.

Add codegen tests, gated to the native path, that assert a real BFloat16
kernel compiles to `bfloat` IR (`fadd bfloat`) and that abs reaches AIR as
the f32-promoted builtin (`air.fabs.f32`) it lacks for bfloat.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
@maleadt

maleadt commented Jun 9, 2026

Copy link
Copy Markdown
Member Author

Turns out that software-emulated BFloat16 already worked. I wonder if we should warn users about this, but it's tricky, since our current bad datatype detection operates at the LLVM level (and software-emulated BFLoat16 is simply i16 there).

@maleadt
maleadt merged commit 960fdcd into main Jun 9, 2026
19 checks passed
@maleadt
maleadt deleted the tb/bfloat16 branch June 9, 2026 17:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

1 participant