Performance¶
This page compares the speed of ace-jax with ML-PACE and MACE, on one CPU and one GPU, for energies and forces.
Each figure shows the throughput against the system size. The throughput is in atom-steps per second (atoms × evaluations per second; higher is better).
- Workload. The workload is similar to MD. Before each timed evaluation of energy and forces, each atom moves a small distance, as between two MD steps. The compilation time is not included.
- Models. The models are medium-size models with random weights.
Their cost is the same as the cost of a fitted model, but they do not
predict correctly:
- PACE
.yacemodels of approximately 500 basis functions for each element (ace-jax and ML-PACE evaluate the same files); - linear ACE models of approximately 450 basis functions for each element;
- MACE-MP-0b2 medium.
- PACE
- Standalone (solid lines) is a direct call of the evaluator, outside an MD engine. For ace-jax and MACE, this is one ASE calculator call from Python, including the neighbour list.
- LAMMPS (dashed lines) is the time of one MD step in LAMMPS: ace-jax
with lammps-jax (LAMMPS export), ML-PACE as
pair_style pace, MACE with Symmetrix. - Systems. SiGe, a random alloy on diamond, and Cantor, the equiatomic CrMnFeCoNi alloy on fcc.
CPU, float64¶

ace-jax runs in LAMMPS only on a GPU, because lammps-jax supports only GPUs. Thus the CPU figure shows ace-jax standalone only.
GPU, float64¶

GPU, float32¶

At 8,192 atoms¶
Atom-steps per second at 8,192 atoms. "—" shows that a code does not run in that setting, or ran out of memory at a smaller size.
| evaluator | mode | CPU, float64: SiGe | CPU, float64: Cantor | GPU, float64: SiGe | GPU, float64: Cantor | GPU, float32: SiGe | GPU, float32: Cantor |
|---|---|---|---|---|---|---|---|
| ace-jax (PACE model) | standalone | 28k | 43k | 892k | 961k | 932k | 1.40M |
| ace-jax (PACE model) | LAMMPS | — | — | 1.09M | 1.33M | 1.79M | 2.04M |
| ace-jax (linear ACE) | standalone | 69k | 73k | 1.47M | 1.48M | 2.04M | 2.03M |
| ace-jax (linear ACE) | LAMMPS | — | — | 2.47M | 1.49M | 3.90M | 2.45M |
| ML-PACE | LAMMPS | 423k | 567k | 2.91M | 3.30M | — | — |
| MACE | standalone | 336 | — | 24k | 19k | 26k | 20k |
| MACE | LAMMPS | 4k | 3k | 60k | 48k | — | — |
| ACEpotentials.jl (linear ACE, direct) | standalone | 95k | 63k | — | — | — | — |
| ACEpotentials.jl (linear ACE, trim library in LAMMPS) | LAMMPS | 1.08M | 944k | — | — | — | — |
Caveats¶
- CPU:
lestrade-cpu, Intel(R) Core(TM) i9-14900K: standalone with 16 threads, LAMMPS with 8 MPI ranks, on CPUs 0-15. - GPU:
sulis-a100, NVIDIA A100-PCIE-40GB. - ACEpotentials.jl standalone is one
energy_forcescall in Julia (neighbour list included), timed after a warm-up call. - ACEpotentials.jl (linear ACE, trim library in LAMMPS) evaluates the exact radial basis; the other ACE lines evaluate splined radials.
- ML-PACE runs only in float64, so it is absent from the float32 figure.
-
MACE in LAMMPS (Symmetrix) evaluates in double whatever the input, so it is shown in float64 only. A MACE line that stops early ran out of memory.
-
Random weights set the cost of a model, not its accuracy. Thus these numbers compare only speed. A fitted model of the same size has the same speed.
- ace-jax standalone uses its neighbour list again while the atoms stay in the skin. MACE standalone builds the list again at each call.
- The throughput depends on the model size, the cutoff and the number of neighbours. The full benchmarks include small and large models, peak memory and the largest system that fits in memory.
More¶
- Full benchmark results: every model size and host, memory, precision, parity checks and versions.
- How to rerun the benchmarks.