Doe Science news source
The DOE Science News Source is a Newswise initiative to promote research news from the Office of Science of the DOE to the public and news media.
  • 2018-03-12 12:30:43
  • Article ID: 690906

A Game Changer: Metagenomic Clustering Powered by HPC

New Berkeley Lab algorithm allows biologists to harness the capabilities of massively parallel supercomputers to make sense of a genomic 'data deluge'

  • Credit: Georgios Pavlopoulos and Nikos Kyrpides, JGI/Berkeley Lab

    Proteins from metagenomes clustered into families according to their taxonomic classification.

  • Credit: Roy Kaltschmidt, Berkeley Lab

    Cori Supercomputer at the National Energy Research Scientific Computing Center (NERSC).

Download HipMCL: https://bitbucket.org/azadcse/hipmcl/

Did you know that the tools used for analyzing relationships between social network users or ranking web pages can also be extremely valuable for making sense of big science data? On a social network like Facebook, each user (person or organization) is represented as a node and the connections (relationships and interactions) between them are called edges. By analyzing these connections, researchers can learn a lot about each user—interests, hobbies, shopping habits, friends, etc.

In biology, similar graph-clustering algorithms can be used to understand the proteins that perform most of life’s functions. It is estimated that the human body alone contains about 100,000 different protein types, and almost all biological tasks—from digestion to immunity—occur when these microorganisms interact with each other. A better understanding of these networks could help researchers determine the effectiveness of a drug or identify potential treatments for a variety of diseases.

Today, advanced high-throughput technologies allow researchers to capture hundreds of millions of proteins, genes and other cellular components at once and in a range of environmental conditions. Clustering algorithms are then applied to these datasets to identify patterns and relationships that may point to structural and functional similarities. Though these techniques have been widely used for more than a decade, they cannot keep up with the torrent of biological data being generated by next-generation sequencers and microarrays. In fact, very few existing algorithms can cluster a biological network containing millions of nodes (proteins) and edges (connections).

That’s why a team of researchers from the Department of Energy’s (DOE’s) Lawrence Berkeley National Laboratory (Berkeley Lab) and Joint Genome Institute (JGI) took one of the most popular clustering approaches in modern biology—the Markov Clustering (MCL) algorithm—and modified it to run quickly, efficiently and at scale on distributed-memory supercomputers. In a test case, their high-performance algorithm—called HipMCL—achieved a previously impossible feat: clustering a large biological network containing about 70 million nodes and 68 billion edges in a couple of hours, using approximately 140,000 processor cores on the National Energy Research Scientific Computing Center’s (NERSC) Cori supercomputer. A paper describing this work was recently published in the journal Nucleic Acids Research.

“The real benefit of HipMCL is its ability to cluster massive biological networks that were impossible to cluster with the existing MCL software, thus allowing us to identify and characterize the novel functional space present in the microbial communities,” says Nikos Kyrpides, who heads JGI’s Microbiome Data Science efforts and the Prokaryote Super Program and is co-author on the paper. “Moreover we can do that without sacrificing any of the sensitivity or accuracy of the original method, which is always the biggest challenge in these sort of scaling efforts.”

“As our data grows, it is becoming even more imperative that we move our tools into high performance computing environments, ” he adds.  “If you were to ask me how big is the protein space? The truth is, we don’t really know because until now we didn’t have the computational tools to effectively cluster all of our genomic data and probe the functional dark matter.” 

In addition to advances in data collection technology, researchers are increasingly opting to share their data in community databases like the Integrated Microbial Genomes & Microbiomes (IMG/M) system, which was developed through a decades-old collaboration between scientists at JGI and Berkeley Lab’s Computational Research Division (CRD). But by allowing users to do comparative analysis and explore the functional capabilities of microbial communities based on their metagenomic sequence, community tools like IMG/M are also contributing to the data explosion in technology.

How Random Walks Lead to Computing Bottlenecks

To get a grip on this torrent of data, researchers rely on cluster analysis, or clustering. This is essentially the task of grouping objects so that items in the same group (cluster) are more similar than those in other clusters. For more than a decade, computational biologists have favored MCL for clustering proteins by similarities and interactions.

“One of the reasons that MCL has been popular among computational biologists is that it is relatively parameter free; users don’t have to set a ton of parameters to get accurate results and it is remarkably stable to small alterations in the data. This is important because you might have to redefine a similarity between data points or you might have to correct for a slight measurement error in your data. In these cases, you don’t want your modifications to change the analysis from 10 clusters to 1,000 clusters,” says Aydin Buluç, a CRD scientist and one of the paper’s co-authors.

But, he adds, the computational biology community is encountering a computing bottleneck because the tool mostly runs on a single computer node, is computationally expensive to execute and has a big memory footprint—all of which limit the amount of data this algorithm can cluster.

One of the most computationally and memory intensive steps in this analysis is a process called random walk. This technique quantifies the strength of a connection between nodes, which is useful for classifying and predicting links in a network. In the case of an Internet search, this may help you find a cheap hotel room in San Francisco for spring break and even tell you the best time to book it. In biology, such a tool could help you identify proteins that are helping your body fight a flu virus.

Given an arbitrary graph or network, it is difficult to know the most efficient way to visit all of the nodes and links. A random walk gets a sense of the footprint by exploring the entire graph randomly; it starts at a node and moves arbitrarily along an edge to a neighboring node. This process keeps going until all of the nodes on the graph network have been reached. Because there are many different ways of traveling between nodes in a network, this step repeats numerous times. Algorithms like MCL will continue running this random walk process until there is no longer a significant difference between the iterations. 

In any given network, you might have a node that is connected to hundreds of nodes and another node with only one connection. The random walks will capture the highly connected nodes because a different path will be detected each time the process is run. With this information, the algorithm can predict with a level of certainty how a node on the network is connected to another. In between each random walk run, the algorithm marks its prediction for each node on the graph in a column of a Markov matrix—kind of like a ledger—and final clusters are revealed at the end. It sounds simple enough, but for protein networks with millions of nodes and billions of edges, this can become an extremely computationally and memory intensive problem. With HipMCL, Berkeley Lab computer scientists used cutting-edge mathematical tools to overcome these limitations. 

“We have notably kept the MCL backbone intact, making HipMCL a massively parallel implementation of the original MCL algorithm,” says Ariful Azad, a computer scientist in CRD and lead author of the paper. 

Although there have been previous attempts to parallelize the MCL algorithm to run on a single GPU, the tool could still only cluster relatively small networks because of memory limitations on a GPU, Azad notes.

“With HipMCL we essentially rework the MCL algorithms to run efficiently, in parallel on thousands of processors, and set it up to take advantage of the aggregate memory available in all compute nodes,” he adds. “The unprecedented scalability of HipMCL comes from its use of state-of-the-art algorithms for sparse matrix manipulation.”

According to Buluç, performing a random walk simultaneously from many nodes of the graph is best computed using sparse-matrix matrix multiplication, which is one of the most basic operations in the recently released GraphBLAS standard. Buluç and Azad developed some of the most scalable parallel algorithms for GraphBLAS’s sparse-matrix matrix multiplication and modified one of their state-of-the-art algorithms for HipMCL.   

“The crux here was to strike the right balance between parallelism and memory consumption. HipMCL dynamically extracts as much parallelism as possible given the available memory allocated to it,” says Buluç.

HipMCL: Clustering at Scale

In addition to the mathematical innovations, another advantage of HipMCL is its ability to run seamlessly on any system—including laptops, workstations and large supercomputers. The researchers achieved this by developing their tools in C++ and using standard MPI and OpenMP libraries.

“We extensively tested HipMCL on Intel Haswell, Ivy Bridge and Knights Landing processors at NERSC, using a up to 2,000 nodes and half a million threads on all processors, and in all of these runs HipMCL successfully clustered networks comprising thousands to billions of edges,” says Buluç. “We see that there is no barrier in the number of processors that it can use to run and find that it can cluster networks 1,000 times faster than the original MCL algorithm.” 

“HipMCL is going to be really transformational for computational biology of big data, just as the IMG and IMG/M systems have been for microbiome genomics,” says Kyrpides. “This accomplishment is a testament to the benefits of interdisciplinary collaboration at Berkeley Lab. As biologists we understand the science, but it’s been so invaluable to be able to collaborate with computer scientists that can help us tackle our limitations and propel us forward.”    

Their next step is to continue to rework HipMCL and other computational biology tools for future exascale systems, which will be able to compute quintillion calculations per second. This will be essential as genomics data continues to grow at a mindboggling rate—doubling about every five to six months.  This will be done as part of DOE Exascale Computing Project’s Exagraph co-design center.

Development of HipMCL was primarily supported by the U.S. Department of Energy’s Office of Science via the  Exascale Solutions for Microbiome Analysis (ExaBiome) project, which is developing exascale algorithms and software to address current limitations in metagenomics research. The development of the fundamental ideas behind this research was also supported by the Office of Advanced Scientific Computing Research’s Applied Math Early Career program. NERSC is an DOE Office of Science User Facility.

HipMCL is based on MPI and OpenMP and is freely available under a modified BSD license at: https://bitbucket.org/azadcse/hipmcl/. Installation instructions and large datasets can be found in the relevant Wiki tab.

X
X
X
  • Filters

  • × Clear Filters

Two Faces Offer Limitless Possibilities

Named for the mythical god with two faces, Janus membranes -- double-sided membranes that serve as gatekeepers between two substances -- have emerged as a material with potential industrial uses.

Relax, Just Break It

Argonne scientists and their collaborators are helping to answer long-held questions about a technologically important class of materials called relaxor ferroelectrics.

Putting Bacteria to Work

Bacteria are diverse and complex creatures that are demonstrating the ability to communicate organism-to-organism and even interact with the moods and perceptions of their hosts (human or otherwise). Scientists call this behavior "bacterial cognition," a systems biology concept that treats these microscopic creatures as beings that can behave like information processing systems.

New Computer Model Predicts How Fracturing Metallic Glass Releases Energy at the Atomic Level

Metallic glasses are an exciting research target for tantalizing applications; however, the difficulties associated with predicting how much energy these materials release when they fracture is slowing down development of metallic glass-based products. Recently, researchers developed a way of simulating to the atomic level how metallic glasses behave as they fracture. This modeling technique could improve computer-aided materials design and help researchers determine the properties of metallic glasses. The duo reports their findings in the Journal of Applied Physics.

The Relationship Between Charge Density Waves and Superconductivity? It's Complicated.

For a long time, physicists have tried to understand the relationship between a periodic pattern of conduction electrons called a charge density wave (CDW), and another quantum order, superconductivity, or zero electrical resistance, in the same material. Do they compete? Co-exist? Co-operate? Do they go their separate ways?

Splitting Water: Nanoscale Imaging Yields Key Insights

In the quest to realize artificial photosynthesis to convert sunlight, water, and carbon dioxide into fuel - just as plants do - researchers need to not only identify materials to efficiently perform photoelectrochemical water splitting, but also to understand why a certain material may or may not work. Now scientists at Lawrence Berkeley National Laboratory have pioneered a technique that uses nanoscale imaging to understand how local, nanoscale properties can affect a material's macroscopic performance.

Feeding Plants to This Algae Could Fuel Your Car

The research shows that a freshwater production strain of microalgae, Auxenochlorella protothecoides, is capable of directly degrading and utilizing non-food plant substrates, such as switchgrass, for improved cell growth and lipid productivity, useful for boosting the algae's potential value as a biofuel.

No More Zigzags: Scientists Uncover Mechanism That Stabilizes Fusion Plasmas

Article describes simulation of physics behind elimination of sawtooth instabilities.

Solutions to Water Challenges Reside at the Interface

Leading Argonne National Laboratory researcher Seth Darling describes the most advanced research innovations that could address global clean water accessibility.

New Cost-Effective Instrument Measures Molecular Dynamics on a Picosecond Timescale

Studying the photochemistry has shown that ultraviolet radiation can set off harmful chemical reactions in the human body and, alternatively, can provide "photo-protection" by dispersing extra energy. To better understand the dynamics of these photochemical processes, a group of scientists irradiated the RNA base uracil with ultraviolet light and documented its behavior on a picosecond timescale. They discuss their work this week in The Journal of Chemical Physics.


  • Filters

  • × Clear Filters

Department of Energy Invests $64 Million in Advanced Nuclear Technology

The U.S. Department of Energy (DOE) has announced nearly $64 million in awards for advanced nuclear energy technology to DOE national laboratories, industry, and 39 U.S. universities in 29 states. Rensselaer Polytechnic Institute has been awarded $800,000 for analysis of nuclear power plants' accident propagation and mitigation processes.

Professor Miao Yu Named the Priti and Mukesh Chatter '82 Career Development Professor

Miao Yu, associate professor in the Howard P. Isermann Department of Chemical and Biological Engineering at Rensselaer Polytechnic Institute, has been named the Priti and Mukesh Chatter Career Development Professor. His research focuses on developing advanced nanomaterials for energy and environmental applications.

Funding for New DOE Energy Frontier Research Center at Brookhaven Lab

UPTON, NY--The U.S. Department of Energy (DOE) has announced funding for a new Energy Frontier Research Center (EFRC) to be led by DOE's Brookhaven National Laboratory. The Brookhaven EFRC, named "Molten Salts in Extreme Environments," will focus on understanding the properties of a class of materials with potential applications in energy technologies--particularly in nuclear power.

Two Stony Brook Researchers Receive Energy Frontier Research Center Awards Totaling $21.75M

Stony Brook University received notification from the U.S. Department of Energy (DOE) that two proposals directed by SBU faculty to expand or develop Energy Frontier Research Centers (EFRCs) designed to accelerate scientific breakthroughs needed to strengthen U.S. economic leadership and energy security will receive funding totaling $21.75 million. The two Stony Brook EFRCs are the Center for Mesoscale Transport Properties (m2M), led by renowned energy storage researcher, Esther Takeuchi, PhD, which will receive a four-year $12 million grant for the existing center; and the creation of a new EFRC, A Next Generation Synthesis Center (GENESIS) led by John Parise, PhD, which will receive a four-year $9.75 million grant.

Seth Davidovits Wins 2018 Marshall N. Rosenbluth Dissertation Award

Article describes dissertation award won by Seth Davidovits.

DOE Launches New Lab Partnering Service

The U.S. Department of Energy officially launched the Lab Partnering Service (LPS), an on-line, single access point platform for investors, innovators, and institutions to identify, locate, and obtain information from DOE's 17 national laboratories.

Department of Energy Announces $75 Million for High Energy Physics Research

The U.S. Department of Energy (DOE) announced $75 million in funding for 77 university research awards on a range of topics in high energy physics to advance knowledge of how the universe works at its most fundamental level.

Thesis Prize Winner's Calculations Characterize Neutrino Interactions

Alessandro Baroni is helping demystify one of the most mysterious particles. His work is contributing to our understanding of neutrinos, and it has earned him the 2017 Jefferson Science Associates Thesis Prize for work performed on a thesis related to research at the Department of Energy's Thomas Jefferson National Accelerator Facility

10 Questions for Steven Cowley, New Director of the Princeton Plasma Physics Laboratory

Steven Cowley, a theoretical physicist and international authority on fusion energy, became the seventh Director of the Princeton Plasma Physics Laboratory (PPon July 1 and will be Princeton professor of astrophysical sciences on September 1.

Ames Laboratory to lead new Center for Advancement of Topological Semimetals

Ames Laboratory will receive $10.75 million over four yearrs for a new Center for Advancement of Topological Semimetals as one of the Department of Energy's Energy Frontier Research Centers.


  • Filters

  • × Clear Filters

Steering Light with Dynamic Lens-on-MEMS

Scientists add active control to design capabilities for new lightweight flat optical devices.

Sugar-Coated Sheets Selectively Target Pathogens

Researchers design self-assembling nanosheets that mimic the surface of cells.

Tracking Down Helium-4's Quarks and Gluons

Scientists obtain the first exclusive measurement of deeply virtual Compton scattering of electrons off helium-4, vital to obtaining an unambiguous 3-D view of quarks and gluons within nuclei.

Predicting Magnetic Explosions: From Plasma Current Sheet Disruption to Fast Magnetic Reconnection

Supercomputer simulations and theoretical analysis shed new light on when and how fast reconnection occurs.

Is Nature Exclusively Left Handed? Using Chilled Atoms to Find Out

Elegant techniques of trapping and polarizing atoms open vistas for beta-decay tests of fundamental symmetries, key to understanding the most basic forces and particles constituting our universe.

As Future Batteries, Hybrid Supercapacitors Are Super-Charged

A new supercapacitor could be a competitive alternative to lithium-ion batteries.

Forever Young Catalyst Reduces Diesel Emissions

Atom probe tomography reveals key explanations for stable performance over a cutting-edge diesel-exhaust catalyst's lifetime.

Sense Like a Shark: Saltwater-Submersible Films

A nickelate thin film senses electric field changes analogous to the electroreception sensing organ in sharks, which detects the bioelectric fields of prey.

A Bit of Quantum Logic--What Did the Atom Say to the Quantum Dot?

Let's talk! Scientists demonstrate coherent coupling between a quantum dot and a donor atom in silicon, vital for moving information inside quantum computers.

New Tech Uses Isomeric Beams to Study How and Where the Galaxy Makes One of Its Most Common Elements

A new measurement using a beam of aluminum-26 prepared in a metastable state allows researchers to better understand the creation of the elements in our galaxy.


Spotlight

Friday July 20, 2018, 03:00 PM

Department of Energy Invests $64 Million in Advanced Nuclear Technology

Rensselaer Polytechnic Institute (RPI)

Thursday July 19, 2018, 05:00 PM

Professor Miao Yu Named the Priti and Mukesh Chatter '82 Career Development Professor

Rensselaer Polytechnic Institute (RPI)

Tuesday July 03, 2018, 11:05 AM

2018 RHIC & AGS Annual Users' Meeting: 'Illuminating the QCD Landscape'

Brookhaven National Laboratory

Friday June 29, 2018, 06:05 PM

Argonne welcomes The Martian author Andy Weir

Argonne National Laboratory

Monday June 18, 2018, 09:55 AM

Creating STEM Knowledge and Innovations to Solve Global Issues Like Water, Food, and Energy

Illinois Mathematics and Science Academy (IMSA)

Friday June 15, 2018, 10:00 AM

Professor Emily Liu Receives $1.8 Million DoE Award for Solar Power Systems Research

Rensselaer Polytechnic Institute (RPI)

Thursday June 07, 2018, 03:05 PM

Celebrating 40 years of empowerment in science

Argonne National Laboratory

Monday May 07, 2018, 10:30 AM

Introducing Graduate Students Across the Globe to Photon Science

Brookhaven National Laboratory

Wednesday May 02, 2018, 04:05 PM

Students from Massachusetts and Washington Win DOE's 28th National Science Bowl(r)

Department of Energy, Office of Science

Thursday April 12, 2018, 07:05 PM

The Race for Young Scientific Minds

Argonne National Laboratory

Wednesday March 14, 2018, 02:05 PM

Q&A: Al Ashley Reflects on His Efforts to Diversify SLAC and Beyond

SLAC National Accelerator Laboratory

Thursday February 15, 2018, 12:05 PM

Insights on Innovation in Energy, Humanitarian Aid Highlight UVA Darden's Net Impact Week

University of Virginia Darden School of Business

Friday February 09, 2018, 11:05 AM

Ivy League Graduate, Writer and Activist with Dyslexia Visits CSUCI to Reframe the Concept of Learning Disabilities

California State University, Channel Islands

Wednesday January 17, 2018, 12:05 PM

Photographer Adam Nadel Selected as Fermilab's New Artist-in-Residence for 2018

Fermi National Accelerator Laboratory (Fermilab)

Wednesday January 17, 2018, 12:05 PM

Fermilab Computing Partners with Argonne, Local Schools for Hour of Code

Fermi National Accelerator Laboratory (Fermilab)

Wednesday December 20, 2017, 01:05 PM

Q&A: Sam Webb Teaches X-Ray Science from a Remote Classroom

SLAC National Accelerator Laboratory

Monday December 18, 2017, 01:05 PM

The Future of Today's Electric Power Systems

Rensselaer Polytechnic Institute (RPI)

Monday December 18, 2017, 12:05 PM

Supporting the Development of Offshore Wind Power Plants

Rensselaer Polytechnic Institute (RPI)

Tuesday October 03, 2017, 01:05 PM

Stairway to Science

Argonne National Laboratory

Thursday September 28, 2017, 12:05 PM

After-School Energy Rush

Argonne National Laboratory

Thursday September 28, 2017, 10:05 AM

Bringing Diversity Into Computational Science Through Student Outreach

Brookhaven National Laboratory

Thursday September 21, 2017, 03:05 PM

From Science to Finance: SLAC Summer Interns Forge New Paths in STEM

SLAC National Accelerator Laboratory

Thursday September 07, 2017, 02:05 PM

Students Discuss 'Cosmic Opportunities' at 45th Annual SLAC Summer Institute

SLAC National Accelerator Laboratory

Thursday August 31, 2017, 05:05 PM

Binghamton University Opens $70 Million Smart Energy Building

Binghamton University, State University of New York

Wednesday August 23, 2017, 05:05 PM

Widening Horizons for High Schoolers with Code

Argonne National Laboratory

Saturday May 20, 2017, 12:05 PM

Rensselaer Polytechnic Institute Graduates Urged to Embrace Change at 211th Commencement

Rensselaer Polytechnic Institute (RPI)

Monday May 15, 2017, 01:05 PM

ORNL, University of Tennessee Launch New Doctoral Program in Data Science

Oak Ridge National Laboratory

Friday April 07, 2017, 11:05 AM

Champions in Science: Profile of Jonathan Kirzner

Department of Energy, Office of Science

Wednesday April 05, 2017, 12:05 PM

High-Schooler Solves College-Level Security Puzzle From Argonne, Sparks Interest in Career

Argonne National Laboratory

Tuesday March 28, 2017, 12:05 PM

Champions in Science: Profile of Jenica Jacobi

Department of Energy, Office of Science

Friday March 24, 2017, 10:40 AM

Great Neck South High School Wins Regional Science Bowl at Brookhaven Lab

Brookhaven National Laboratory

Wednesday February 15, 2017, 04:05 PM

Middle Schoolers Test Their Knowledge at Science Bowl Competition

Argonne National Laboratory

Friday January 27, 2017, 04:00 PM

Haslam Visits ORNL to Highlight State's Role in Discovering Tennessine

Oak Ridge National Laboratory

Tuesday November 08, 2016, 12:05 PM

Internship Program Helps Foster Development of Future Nuclear Scientists

Oak Ridge National Laboratory

Friday May 13, 2016, 04:05 PM

More Than 12,000 Explore Jefferson Lab During April 30 Open House

Thomas Jefferson National Accelerator Facility

Monday April 25, 2016, 05:05 PM

Giving Back to National Science Bowl

Ames Laboratory

Friday March 25, 2016, 12:05 PM

NMSU Undergrad Tackles 3D Particle Scattering Animations After Receiving JSA Research Assistantship

Thomas Jefferson National Accelerator Facility

Tuesday February 02, 2016, 10:05 AM

Shannon Greco: A Self-Described "STEM Education Zealot"

Princeton Plasma Physics Laboratory

Monday November 16, 2015, 04:05 PM

Rare Earths for Life: An 85th Birthday Visit with Mr. Rare Earth

Ames Laboratory

Tuesday October 20, 2015, 01:05 PM

Meet Robert Palomino: 'Give Everything a Shot!'

Brookhaven National Laboratory

Tuesday April 22, 2014, 11:30 AM

University of Utah Makes Solar Accessible

University of Utah

Wednesday March 06, 2013, 03:40 PM

Student Innovator at Rensselaer Polytechnic Institute Seeks Brighter, Smarter, and More Efficient LEDs

Rensselaer Polytechnic Institute (RPI)

Friday November 16, 2012, 10:00 AM

Texas Tech Energy Commerce Students, Community Light up Tent City

Texas Tech University

Wednesday November 23, 2011, 10:45 AM

Don't Get 'Frosted' Over Heating Your Home This Winter

Temple University

Wednesday July 06, 2011, 06:00 PM

New Research Center To Tackle Critical Challenges Related to Aircraft Design, Wind Energy, Smart Buildings

Rensselaer Polytechnic Institute (RPI)

Friday April 22, 2011, 09:00 AM

First Polymer Solar-Thermal Device Heats Home, Saves Money

Wake Forest University

Friday April 15, 2011, 12:25 PM

Like Superman, American University Will Get Its Energy from the Sun

American University





Showing results

0-4 Of 2215