89 papers, 2 posters and 2 further items published between 2021 and 2025, listed by year, most recent first. Some links point to preprints where the published version sits behind a paywall.
2025
- Optimizing and Scaling Machine Learning Models for Scientific Applications on Exascale Supercomputers.
- Closing a Source Complexity Gap between Chapel and HPX.
- Generalized Ideal Point Models for Robust Measurement with Dirty Data in the Social Sciences.
- The AmpereOne A192-32X in Perspective: Benchmarking a New Standard.
- Towards Automatic, Predictable and High-Performance Parallel Code Generation.
- Metagenomes and Metagenome-Assembled Genomes from Tidal Lagoons at a New York City Waterfront Park.
- ByteBoost: An advanced cybertraining program designed to enhance research on testbed systems.
- Continuous HPC Performance Monitoring: Can It Run Without Affecting User Jobs?.
- Enhancing Tensor Program Tuning With Transfer Learning and Hardware-Aware Polyhedral Optimizations.
2024
- Cross-Feature Transfer Learning For Efficient Tensor Program Generation.
- Impact of Write-Allocate Elimination on Fujitsu A64FX.
- First Impressions of the NVIDIA Grace CPU Superchip and NVIDIA Grace Hopper Superchip for Scientific Workloads.
- Parallel C++ Efficient and Scalable High-Performance Parallel Programming Using HPX.
- Quantifying potential marine debris sources and potential threats to penguins on the West Antarctic Peninsula.
- Anti-Coulomb ion-ion interactions: A theoretical and computational study.
- Parallel assembly of finite element matrices on multicore computers.
- First Impressions of the Sapphire Rapids Processor with HBM for Scientific Workloads.
- Performance-Portable Tensor Transpositions in MLIR.
- A64FX Enables Engine Decarbonization Using Deep Learning.
- From array expressions to predictable portable high-performance: foundations for no-code HPC on arrays.
- Accelerating LULESH using HPX – the C++ Standard Library for Parallelism and Concurrency.
- Hardware-Software Co-design of Efficient and Scalable Deep Learning.
- Dynamics of Jet Expansion and Impingement Across a Spectrum of Nozzle Pressure Ratios.
- Benchmarking with Supernovae: A Performance Study of the FLASH Code.
- Exploring Processor Micro-architectures Optimised for BLAS3 Micro-kernels.
- Enhancing Code Portability, Problem Scale, and Storage Efficiency in Exascale Applications.
- Towards a Scalable and Efficient PGAS-based Distributed OpenMP.
- On the Scalability of Computing Genomic Diversity Using SparkLeBLAST: A Feasibility Study.
- Benchmarking and Continuous Performance Monitoring of Ookami, an ARM Fujitsu A64FX Testbed Cluster.
- From Saline to Solids: Studies of Ionic Solvation and Machine Learning for Ab Initio Calculations.
- Xphase3d: Memory-Distributed Phase Retrieval for Reconstructing Large-Scale 3D Density Maps of Biological Macromolecules.
- Improving Polyhedral-Based Optimizations With Dynamic Coordinate Descent.
- Evaluating Tuning Opportunities of the LLVM/OpenMP Runtime.
- Studying CPU and memory utilization of applications on Fujitsu A64FX and Nvidia Grace Superchip.
- Asynchronous-Many-Task Systems: Challenges and Opportunities -- Scaling an AMR Astrophysics Code on Exascale machines using Kokkos and HPX.
- Preparing for HPC on RISC-V: Examining Vectorization and Distributed Performance of an Astrophyiscs Application with HPX and Kokkos.
- Simulating Stellar Merger using HPX/Kokkos on A64FX on Supercomputer Fugaku.
- Explore as a Storm, Exploit as a Raindrop: On the Benefit of Fine-Tuning Kernel Schedulers with Coordinate Descent.
2023
- OpenMP Advisor: A Compiler Tool for Heterogenous Architectures.
- Examining the Connectivity of Antarctic Krill on the West Antarctic Peninsula: Implications for Pygoscelis Penguin Biogeography and Population Dynamics.
- Are we ready for broader adoption of ARM in the HPC community: Performance and Energy Efficiency Analysis of Benchmarks and Applications Executed on High-End ARM Systems.
- Performance Study on CPU-based Machine Learning with PyTorch.
- Shared memory parallelism in Modern C++ and HPX.
- Interoperable PGAS Programming Models for Exascale Supercomputing.
- Program Transformation for Automatic GPU-Offloading using OpenMP.
- HEAT: A Highly Efficient and Affordable Training System for Collaborative Filtering Based Recommendation on CPUs.
- Asynchronous Many-Task Systems and Applications: First International Workshop.
- CPU Architecture Modelling and Design.
- Cyberinfrastructure for Sustainability Sciences.
- Human mobility patterns are associated with experienced partisan segregation in US metropolitan areas.
- Quantifying Antarctic krill connectivity across the West Antarctic Peninsula and its role in large-scale Pygoscelis penguin population dynamics.
- Sinophobia was popular in Chinese language communities on Twitter during the early COVID-19 pandemic.
- From Molecular Dynamics to Oceanography - Ookami Graduate Students Porting and Tuning Science Codes for A64FX.
- A Further Study of Linux Kernel Hugepages on A64FX with FLASH, an Astrophysical Simulation Code.
- LM4HPC: Towards Effective Language Model Application in High-Performance Computing.
- Evaluating HPX and Kokkos on RISC-V using an Astrophysics Application Octo-Tiger.
- Efficient Auto-Vectorization for Control-flow Dependent Loops through Data Permutation.
- Parameterization of Quantum Interactions.
- Ookami: An A64FX Computing Resource.
- Benchmarking the Parallel 1D Heat Equation Solver in Chapel, Charm++, C++, HPX, Go, Julia, Python, Rust, Swift, and Java.
- The General Atomic and Molecular Electronic Structure System (GAMESS): Novel Methods on Novel Architectures.
2022
- Experiences with Porting the FLASH Code to Ookami, an HPE Apollo 80 A64FX Platform.
- OpenSHMEM Active Message Extension for Task-Based Programming.
- Analysis of Vector Particle-In-Cell (VPIC) memory usage optimizations on cutting-edge computer architectures.
- Dirac lines and loop at the Fermi level in the time-reversal symmetry breaking superconductor LaNiGa2.
- Parthenon – a performance portable block-structured adaptive mesh refinement framework.
- Quantum calcium-ion affective influences measured by EEG.
- Exploring Source-to-Source Compiler Transformation of OpenMP SIMD Constructs for Intel AVX and Arm SVE Vector Architectures.
- FOURST: A code generator for FFT-based fast stencil computations.
- Friends and foes: Sinophobia was viral in Chinese language communities on Twitter during the early COVID-19 pandemic.
- Developing Accurate Slurm Simulator.
- On Using Linux Kernel Huge Pages with FLASH, an Astrophysical Simulation Code.
- Performance of an Astrophysical Radiation Hydrodynamics Code under Scalable Vector Extension Optimization.
- Bring the BitCODE - Moving Compute and Data in Distributed Heterogeneous Systems.
- From Merging Frameworks to Merging Stars: Experiences using HPX, Kokkos and SIMD Types.
- Assessing the State of Autovectorization Support based on SVE.
- Improved Distributed-memory Triangle Counting by Exploiting the Graph Structure.
- Modern server ARM processors for supercomputers: A64FX and others. Initial data of benchmarks.
- Vectorizing divergent control flow with active-lane consolidation on long-vector architectures.
2021
- VPIC 2.0: Next Generation Particle-in-Cell Simulations.
- Ookami: Deployment and Initial Experiences.
- Comparing the behavior of OpenMP Implementations with various Applications on two different Fujitsu A64FX platforms.
- MoB2 under Pressure: Superconducting Mo Enhanced by Boron.
- A64FX performance: experience on Ookami.
- Porting and Evaluation of a Distributed Task-driven Stencil-based Application.
- Comparing OpenMP Implementations with Applications Across A64FX Platforms.
- Educating HPC users in the use of advanced computing technology.
- Towards Architecture-aware Hierarchical Communication Trees on Modern HPC Systems.
- Hybrid Classical-Quantum Computing: Applications to Statistical Mechanics of Neocortical Interactions.
Posters
- Stalls and Memory Analysis on Fujitsu A64FX and NVIDIA Grace.
- Pure Deflagrations of Hybrid CONe White Dwarf Progenitors.
Other
- Kernel module for the A64FX hardware barrier.
- PEARC 2022 - Birds of a feather session: NSF innovative computing technology testbed community exchange.
