Skip to content

Performance

This page compares the speed of ace-jax with ML-PACE and MACE, on one CPU and one GPU, for energies and forces.

Each figure shows the throughput against the system size. The throughput is in atom-steps per second (atoms × evaluations per second; higher is better).

  • Workload. The workload is similar to MD. Before each timed evaluation of energy and forces, each atom moves a small distance, as between two MD steps. The compilation time is not included.
  • Models. The models are medium-size models with random weights. Their cost is the same as the cost of a fitted model, but they do not predict correctly:
    • PACE .yace models of approximately 500 basis functions for each element (ace-jax and ML-PACE evaluate the same files);
    • linear ACE models of approximately 450 basis functions for each element;
    • MACE-MP-0b2 medium.
  • Standalone (solid lines) is a direct call of the evaluator, outside an MD engine. For ace-jax and MACE, this is one ASE calculator call from Python, including the neighbour list.
  • LAMMPS (dashed lines) is the time of one MD step in LAMMPS: ace-jax with lammps-jax (LAMMPS export), ML-PACE as pair_style pace, MACE with Symmetrix.
  • Systems. SiGe, a random alloy on diamond, and Cantor, the equiatomic CrMnFeCoNi alloy on fcc.

CPU, float64

Throughput against system size on CPU in float64, SiGe and Cantor panels: ACEpotentials.jl trim library and ML-PACE in LAMMPS are fastest, ace-jax standalone 6–10× below ML-PACE and level with ACEpotentials.jl direct, MACE slowest

ace-jax runs in LAMMPS only on a GPU, because lammps-jax supports only GPUs. Thus the CPU figure shows ace-jax standalone only.

GPU, float64

Throughput against system size on GPU in float64, SiGe and Cantor panels, for ace-jax PACE and linear ACE, ML-PACE and MACE, in LAMMPS and standalone

GPU, float32

Throughput against system size on GPU in float32, SiGe and Cantor panels, for ace-jax PACE and linear ACE and MACE standalone

At 8,192 atoms

Atom-steps per second at 8,192 atoms. "—" shows that a code does not run in that setting, or ran out of memory at a smaller size.

evaluator mode CPU, float64: SiGe CPU, float64: Cantor GPU, float64: SiGe GPU, float64: Cantor GPU, float32: SiGe GPU, float32: Cantor
ace-jax (PACE model) standalone 28k 43k 892k 961k 932k 1.40M
ace-jax (PACE model) LAMMPS — — 1.09M 1.33M 1.79M 2.04M
ace-jax (linear ACE) standalone 69k 73k 1.47M 1.48M 2.04M 2.03M
ace-jax (linear ACE) LAMMPS — — 2.47M 1.49M 3.90M 2.45M
ML-PACE LAMMPS 423k 567k 2.91M 3.30M — —
MACE standalone 336 — 24k 19k 26k 20k
MACE LAMMPS 4k 3k 60k 48k — —
ACEpotentials.jl (linear ACE, direct) standalone 95k 63k — — — —
ACEpotentials.jl (linear ACE, trim library in LAMMPS) LAMMPS 1.08M 944k — — — —

Caveats

  • CPU: lestrade-cpu, Intel(R) Core(TM) i9-14900K: standalone with 16 threads, LAMMPS with 8 MPI ranks, on CPUs 0-15.
  • GPU: sulis-a100, NVIDIA A100-PCIE-40GB.
  • ACEpotentials.jl standalone is one energy_forces call in Julia (neighbour list included), timed after a warm-up call.
  • ACEpotentials.jl (linear ACE, trim library in LAMMPS) evaluates the exact radial basis; the other ACE lines evaluate splined radials.
  • ML-PACE runs only in float64, so it is absent from the float32 figure.
  • MACE in LAMMPS (Symmetrix) evaluates in double whatever the input, so it is shown in float64 only. A MACE line that stops early ran out of memory.

  • Random weights set the cost of a model, not its accuracy. Thus these numbers compare only speed. A fitted model of the same size has the same speed.

  • ace-jax standalone uses its neighbour list again while the atoms stay in the skin. MACE standalone builds the list again at each call.
  • The throughput depends on the model size, the cutoff and the number of neighbours. The full benchmarks include small and large models, peak memory and the largest system that fits in memory.

More