The NIH Aging and AD/ADRD Omics Data Resources: A Path to Interoperability

Image
Purple outline of a human brain over a purple background

Audience

Data scientists, informaticians, ethicists, Alzheimer's Disease and Alzheimer's Disease Related Dementias (AD/ADRD) researchers, and research participants with an interest in establishing FAIR (Findable, Accessible, Interoperable, and Reusable) and interoperable AD/ADRD resources.

Dates

July 11, 2024 | 8:30 a.m.–4:30 p.m. ET
July 12, 2024 | 8:30 a.m.–3:00 p.m. ET

Purpose and Background

On July 11-12, 2024, NIA hosted a hybrid workshop convening AD/ADRD omics data and related resources across and outside the NIH to explore opportunities and challenges for establishing an equitable, effective omics data ecosystem to accelerate and improve NIA-supported research.

This workshop built on the NIH AD/ADRD Platforms Workshop, FAIRness Within and Across Data Infrastructures , conducted in June 2023 and continued the discussion by convening experts from diverse fields to address current challenges and opportunities in data interoperability, representation, and usage for both Aging and AD/ADRD data resources. A set of stories from the perspective of end users illustrated current challenges and opportunities from the following perspectives and provide a framework for the workshop.

  • Diversity, Equity, and Inclusivity in Omics Research
  • Global Data Sharing: Current Challenges and Future Opportunities
  • Omics Research in a Distributed Environment: Computational Challenges and Solutions
  • Creating Interoperability Between Model System and Human Omics Data in Aging and AD/ADRD

Agenda

See below for agenda information for each day.

Day 1 | July 11, 2024

8:30 a.m. Welcoming Remarks, Amy S. Kelley, M.D., M.S.H.S, and Michael Bennani, Ph.D., NIA

8:50 a.m. Workshop Background
Co-Chairs: Lucila Ohno-Machado, M.D., Ph.D., M.B.A, Yale University and David Bennett, M.D., Rush University

  • User Stories Session 2: Sid O'Bryant, Ph.D., University of North Texas Health Science Center
  • User Stories Session 3: Heidi Sofia, Ph.D., National Center for Biotechnology Information (NCBI)
  • User Stories Session 4: Brian O'Connor, Ph.D., Nimbus Informatics, LLC
  • User Stories Session 5: Melissa Haendel, Ph.D., FACMI; University of North Carolina

9:35 a.m. Session 1 | Introductions to Aging and AD/ADRD Data Platforms and Other Data Ecosystems
Co-Chairs: Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

  • National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (NIAGADS), Li-San Wang, Ph.D., University of Pennsylvania
  • AD Knowledge Portal and Exceptional Longevity Translational Resources (ELITE) Portal, Anna Greenwood, Ph.D., Sage Bionetworks
  • Alzheimer's Disease Data Initiative (ADDI), Matt Clement, Ph.D., Alzheimer’s Disease Data Initiative and Mukta Phatak, Ph.D., Gates Ventures

10:10 a.m. Morning Break

  • National Center for Biotechnology Information (NCBI) database of Genotypes and Phenotypes (dbGaP), Kim Pruitt, Ph.D., NCBI
  • National Heart, Lung, and Blood Institute (NHLBI) BioData Catalyst, Regina Bures, Ph.D., NHLBI
  • National Human Genome Research Institute (NHGRI) Analysis, Visualization and Informatics Lab-space (AnVIL), Chris Wellington, BSc, NHGRI
  • US Dept. of Veteran Affairs’ (VA) Million Veteran Program (MVP), Mark Logue, Ph.D., VA
  • Rapid Acceleration of Diagnostics (RADX-TDR), Matt Anderson, Ph.D., University of Wisconsin-Madison
  • Audience Q & A, Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

11:40 a.m. Lunch Break

12:40 p.m. Session 2 | Diversity, Equity, and Inclusivity in Omics Research
Chair: Sid O'Bryant, Ph.D., University of North Texas Health Science Center

  • Genomics through the lens of lived experiences and environmental influences affecting communities with varied ancestral backgrounds, Robert Barber, Ph.D., University of North Texas Health Science Center
  • Assessment of redlining practices that impact health outcomes in highly vulnerable neighborhoods, Data analytics, algorithmic biases, and research integrity with emphasis on social vulnerability and the integration of genomics with non-genomics data to shape scientific knowledge, Latrice Landry, Ph.D., University of Pennsylvania
  • Understanding the role of APOE-E4, social determinants of health, and comorbidities in AT(N)-V imaging markers in diverse populations, Karin Meeker, Ph.D.; University of North Texas Health Science Center
  • Community values and efforts to communicate research findings from omics studies to research participant communities, Barbara Harrison, MS, CGC, Howard University
  • User experience challenges facing diverse omics data end-users and its impact on research while addressing viable solutions that may increase data access, Krystal Tsosie, Ph.D., M.P.H., M.A., Arizona State University
  • Data access and data use challenges that are impacted by international and national regulations and laws, All Session 5 Speakers/Chair
  • Audience Q & A, Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

2:40 p.m. Afternoon Break

2:55 p.m. Session 3 | Global Data Sharing, Current Challenges and Future Opportunities
Chair: Heidi Sofia, Ph.D., National Center for Biotechnology Information (NCBI)

  • Enhancing Precision Medicine in Latin America: Ancestry, Admixture, and Genomic Data Sharing, Iscia Lopes-Cendes, M.D., Ph.D., University of Campinas, Brazil
  • Federated Data Access and Computing in Africa, Takudzwa Nyasha Musarurwa, M.S., University of Cape Town, South Africa
  • Privacy-preserving federated analytics in Swiss hospitals and beyond, Jean-Pierre Hubaux, Dr.Eng., École polytechnique fédérale de Lausanne (EPFL), Switzerland
  • Encouraging open data provision: challenges and opportunities, Matteo Tranchero, M.S., University of Pennsylvania
  • Audience Q & A, Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

4:30 p.m. Adjourn, Mette Peters, Ph.D., NIA

Day 2 | July 12, 2024

8:30 a.m. Summary of Day 1, Lucila Ohno-Machado, M.D., Ph.D., M.B.A., Yale University and David Bennett, M.D., Rush University

8:45 a.m. Session 4 | Omics Research in a Distributed Environment, Computational Challenges and Solutions
Chair: Brian O'Connor, Ph.D., Nimbus Informatics, LLC

  • The NIH Cloud Platform Interoperability efforts, Valentina Di Francesco, M.S., NHGRI
  • Advanced Research Projects Agency for Health (ARPA-H) BDF Toolbox, Erika Kim, Ph.D., ARPA-H
  • Compute beyond the shining sea, Jack DiGiovanna, Ph.D., Velsera
  • The National Alzheimer’s Coordinating Center (NACC)’s Data Ecosystem, Revolutionizing Multimodal Data Integration, Interoperability, and Access to Advance Alzheimer’s Disease and Related Dementia Discovery and Translation, Sarah Biber, Ph.D., NACC
  • Leveraging AI to harmonize data at scale, Shannon Ballard, Ph.D., Intramural Center for Alzheimer’s and Related Dementias (CARD)
  • AI, data analytics, and the integration of environmental data into -omic datasets, Chirag Patel, Ph.D., Harvard Medical School
  • Audience Q & A, Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

10:45 a.m. Morning Break

11:00 a.m. Session 5 | Creating Interoperability Between Model System and Human Omics Data in Aging and AD/ADRD
Chair: Melissa Haendel, Ph.D., FACMI, University of North Carolina

  • Translational data analysis for ADRD in the MODEL-AD consortium, Gregory W. Carter, Ph.D., The Jackson Laboratory
  • The Monarch Initiative, Leveraging model organisms to characterize phenotype to disease associations in AD, Monica Munoz-Torres, Ph.D., University of Colorado
  • Supporting LLMs with Structured Relationships to Integrate Knowledge of Models of Aging and AD, Harry Caufield, Ph.D., Lawrence Berkeley National Laboratory
  • Human iPSC-based experimental systems for capturing genetic diversity and risk for AD/ADRDs, Tracy Young-Pearse, Ph.D., Harvard Medical School
  • Audience Q & A, Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

12:30 p.m. Lunch Break

1:15 p.m. Workshop Summary

  • Panel Discussion with Session Chairs, Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA
  • Next Steps, Lucila Ohno-Machado, M.D., Ph.D., MBA, Yale University and David Bennett, M.D., Rush University

2:30 p.m. Closing Remarks, Jennie Larkin, Ph.D., NIA

Executive Summary

The NIH Aging and AD/ADRD Omics Data Resources: A Path to Interoperability workshop was held on July 11-12, 2024. This summary highlights findings and conclusions for each of the discussions.

Workshop materials are also being shared through the Open Science Foundation (OSF), which allows the content to be linked through a Digital Object Identifier (DOI), promotes findability, accessibility, interoperability, and reusability (FAIR) practices, and provides citation instructions.

Executive Summary

The National Institute on Aging (NIA) supports valuable infrastructures that provide data resources to enable and accelerate research on aging, Alzheimer’s disease, and Alzheimer’s disease-related dementias (AD/ADRD). The National Institutes of Health (NIH) Aging and AD/ADRD Omics Data Resources: A Path to Interoperability workshop convened aging and AD/ADRD omics data and related resources, as well as key stakeholders and subject matter experts across and outside NIH to explore opportunities and challenges for establishing an equitable, effective omics data ecosystem to accelerate and improve NIA-supported research. This workshop built on the June 2023 NIH AD/ADRD Platforms Workshop: FAIRness (i.e., findability, accessibility, interoperability, and reusability) Within and Across Data Infrastructures , and continued the discussion by convening key stakeholders from diverse fields to address current challenges, opportunities, and solutions relevant to data interoperability, representation, and usage for both aging and AD/ADRD data resources.

The two-day workshop focused on stories from the perspective of data end users that illustrated current challenges and opportunities from various perspectives, including (1) diversity, equity, and inclusivity in omics research; (2) global data sharing; (3) omics research in a distributed environment; and (4) creating interoperability between model system and human omics data in aging and AD/ADRD.

Welcoming Remarks and Workshop Background

Dr. Amy Kelley welcomed participants and explained NIA’s interest in exploring opportunities to improve data interoperability, which will provide valuable data infrastructure, support data accessibility, and accelerate aging and AD/ADRD research. The objectives of this workshop are to discuss the aging and AD/ADRD data ecosystem and consider what is necessary to establish an equitable and effective omics data ecosystem that will accelerate research discoveries. The input from this workshop will be integral to informing future strategic efforts in aging and AD/ADRD omics research. Dr. Kelley was followed by workshop co-chair Dr. Lucila Ohno-Machado who reviewed the key themes of the workshop.

To set the stage, the chairs of sessions 2–5 provided overviews and shared user stories related to the session topic. Each session overview highlighted challenges and opportunities for aging and AD/ADRD omics research, including how to promote diversity, equity, and inclusivity; enhance global data sharing; address difficulties with distributed research environments; and create interoperability between model systems and human data.

Session 1: Introductions to Aging and AD/ADRD Data Platforms and Other Data Ecosystems

Drs. Michael Bennani and Mette Peters moderated presentations and discussions related to aging and AD/ADRD data platforms and ecosystems. Dr. Li-San Wang presented on the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site ( NIAGADS ), the genetic data repository and data coordinating center for the Alzheimer’s Disease Sequencing Project ( ADSP ). Dr. Anna Greenwood provided an overview of the AD Knowledge Portal and the Exceptional Longevity Translational Resources ( ELITE ) Portal. Both portals contain human and model system omics and other data types. Drs. Matt Clement and Mukta Phatak described the Alzheimer's Disease Data Initiative ( ADDI ) and ADDI’s AD Workbench , which enables federated access to datasets. Dr. Kim Pruitt presented on the National Center for Biotechnology Information (NCBI) database of Genotypes and Phenotypes ( dbGaP ), an archive of DNA genotypes and phenotypes from NIH-funded studies. Dr. Regina Bures from the National Heart, Lung, and Blood Institute (NHLBI) shared that BioData Catalyst ( BDC ) contains multimodal data and is interoperable with several federal repositories. Dr. Chris Wellington discussed the Analysis, Visualization and Informatics Lab-space ( AnVIL ) platform, a federate data storage, analysis and sharing platform with genomic, phenotypic, and molecular data. Dr. Mark Logue from the US Department of Veterans Affairs (VA) provided an overview of the Million Veteran Program ( MVP ). MVP is only accessible to VA-funded investigators with MVP approved projects. Dr. Matt Anderson discussed the D4I Tribal Data Repository , which houses American Indian/Alaska Native data from the Rapid Acceleration of Diagnostics Underserved Populations ( RADx-UP ) program. Common themes from the presentations and discussion included developing universal interoperability standards, harmonizing data access processes and policies, and incentivizing researchers to use cloud-based platforms for access and compute to enhance data security.

Session 2: Diversity, Equity, and Inclusivity in Omics Research

Dr. Sid O’Bryant moderated presentations and discussion for the session. He highlighted opportunities and challenges for gathering data, accessing knowledge and resources, publishing and presenting findings, sharing findings with the broader community, and educating the scientific community. Dr. Robert Barber highlighted the evidence that race and ethnicity are artificial rather than biological constructs and the need to focus on how factors like social determinants of health (SDOH) drive AD/ADRD disease onset and progression. Dr. Latrice Landry emphasized the need to study how aspects of a person’s lived experience, like redlining and social vulnerability, need to be incorporated into genomic and health disparities research. Ms. Barbara Harrison noted examples of scientific research that harmed individuals and shared the steps required for successful community-based participatory research (CBPR). Key CBPR principles were discussed and included dissemination, cultural competency, transparency, and capacity. Dr. Krystal Tsosie noted that Indigenous people, including American Indian and Alaska Native people, are often mistrustful of research and that informed consent processes and messaging are frequently inconsistent with Indigenous culture or misleading. She provided suggestions for addressing these issues, including informed consent, data governance, and data access standards that considers Indigenous people. Some topics that emerged from the presentations and discussion included ways to engage with diverse, marginalized, and underrepresented communities and involve them in omics research, such as conversations between these communities and omics researchers and including community members on data access committees. There were also discussions about potential solutions to encourage researchers to increase diversity and representation in their research, such as infrastructure grants to support community engagement in research projects. The presenters also discussed barriers to achieving diversity, equity, and inclusivity in aging and AD/ADRD omics research, including funding and lack of awareness on the importance of increasing representation and how to effectively engage underrepresented communities.

Session 3: Global Data Sharing, Current Challenges, and Future Opportunities

Dr. Heidi Sofia moderated presentations and discussion related to the current challenges and future opportunities for global data sharing. As part of her introduction of the session, Dr. Sofia highlighted the Global Alliances for Genomics and Health ( GA4GH ), an international organization dedicated to developing data sharing standards and policies that ensure everyone experiences the benefits of scientific advancement. Dr. Iscia Lopes-Cendes said that the Brazilian Initiative of Precision Medicine ( BIPMed ) is studying the admixture in the Brazilian population to support genomic research, implement precision medicine, and improve genetic testing in Brazil. Mr. Takudzwa Nyasha Musarurwa from the Data Science for Health Discovery and Innovation in Africa ( DS-I Africa ) program provided an overview of the work to improve data sharing and compute interoperability on the eLwazi platform using the GA4GH starter kit service. Dr. Jean-Pierre Hubaux summarized privacy enhancing technologies (PETs) that support data sharing without the data leaving its protected environment and highlighted the applications of Tune Insight , a privacy-preserving federated artificial intelligence (AI) tool. Mr. Matteo Tranchero noted the lack of incentives for researchers to share their data and proposed using entitymetrics as a credit model for scientific ideas and innovation that does not place weight on the number of citations of a publication. The discussion focused on using the GA4GH starter kit for other platforms, the pros and cons of entitymetrics, and challenges with technological infrastructures adopting PETs.

Session 4: Omics Research in a Distributed Environment, Computational Challenges, and Solutions

Dr. Brian O’Connor moderated presentations and discussion for the session, but first highlighted challenges researchers experience with finding and working and computing with multimodal data on distributed, disparate systems. Ms. Valentina Di Francesco shared the work of the NIH Cloud Platform Interoperability ( NCPI ) Program to develop and implement standards that enable interoperability and facilitate an NIH federated data ecosystem. Dr. Erika Kim explained the structure, research goals, and projects of the Advanced Research Projects Agency for Health (ARPA-H) Biomedical Data Fabric ( BDF ) Toolbox program. Dr. Jack DiGiovanna shared examples of successful global data sharing, which require coordinated technology and policy efforts. Dr. Sarah Biber explained the infrastructure, data, tools, and resources of the National Alzheimer’s Coordinating Center ( NACC ) that support interoperability among data from the Alzheimer’s Disease Research Centers ( ADRCs ). Dr. Shannon Ballard and Mr. Alan Long discussed the development of the Data Inventory and Validation Environment for Research (DIVER) and Generative Common Data Elements (GenCDE) tools, which are AI tools that support data harmonization. Dr. Chirag Patel highlighted the challenges of conducting exposomics research across data platforms and proposed multimodal AI and informatics methods to standardize how these data are collected and shared. The discussion focused on ways to improve interoperability and collect and harmonize data for more impactful exposomic research.

Session 5: Creating Interoperability Between Model System and Human Omics Data in Aging and AD/ADRD

Dr. Melissa Haendel moderated the presentations and discussion and noted that the lack of interoperability between aging and AD/ADRD human studies and model systems has limited research focused on characterizing and treating AD/ADRD. Dr. Gregory Carter shared the work of the Model Organism Development and Evaluation for Late-onset Alzheimer’s Disease ( MODEL-AD ) consortium to develop AD/ADRD mouse models that reflect the complex etiology of late-onset AD and can be used for therapeutic development. Dr. Monica Munoz-Torres discussed the Monarch Initiative resource, which brings together cross-species, multimodal, disparate data by unifying ontologies and leveraging semantic data models. Dr. Harry Caufield demonstrated how large language models (LLMs) can be used to retrieve information from literature and understand study methods and observations made in different model systems. Dr. Tracy Young-Pearse shared her work to characterize AD/ADRD using induced pluripotent stem cells (iPSCs). The discussion covered how to address the “Aim 3 phenomenon,” which is the issue of only including modeling studies, computational resources, and interoperability resources in Aim 3 rather than integrating these details throughout a project. There was also a discussion about creating data standards for model systems.

Workshop Summary

Based on each of the session themes, the session chairs shared actionable items that were designated by importance (higher or lower) and implementation timeline (quick or longer). The actionable items for session 2 (Diversity, Equity, and Inclusivity in Omics Research) included:

  • Higher Importance/Quick Implementation: Establish a resource for investigators to identify experts and datasets to aid in the advancements of omics research among diverse communities.
  • Lower Importance/Quick Implementation: Establish a resource for investigators to identify resources to be able to conduct scientific analyses for the advancement of omic research among diverse communities.
  • Higher Importance/Longer Implementation: Establish a recruitment resource for existing NIA- funded programs to increase representation in clinical research.
  • Lower Importance/Longer Implementation: Establish a resource for modeling novel findings from diverse and representative communities in nonhuman systems.

The actionable items for session 3 (Global Data Sharing) were:

  • Higher Importance/Quick Implementation: Drive broad adoption of GA4GH standards in technical platforms and policy frameworks.
  • Lower Importance/Quick Implementation: Contribute genomic data to the global pangenome reference.
  • Higher Importance/Longer Implementation: Implement privacy-enhancing technologies that are transparent to the user into data resources, and use entitymetrics to measure data value, improve credit models, and promote fair models of benefits.
  • Lower Importance/Longer Implementation: Link global data resources in an interoperable
    network using ARPA-H data services.

The actionable items for session 4 (Omics Research in a Distributed Environment) support technical interoperability standards and included:

  • Higher Importance/Quick Implementation: Implement RAS passports.
  • Lower Importance/Quick Implementation: Adopt DRS1.5.
  • Higher Importance/Longer Implementation: Establish federated compute systems.
  • Lower Importance/Longer Implementation: Create a data catalog.

The actionable items for session 5 (Creating Interoperability Between Model System and Human Omics Data in Aging and AD/ADRD were:

  • Higher Importance/Quick Implementation: Design model system experiments to match real- world patient populations.
  • Lower Importance/Quick Implementation: Require AD programs to include collaboration with experts in standards and interoperability, especially phenotyping and biomarkers.
  • Higher Importance/Longer Implementation: Increase focus on standardizing environmental and social measures in patients and model systems and utilize natural, diverse, and richly longitudinal cohorts together with real-world and model system data.
  • Lower Importance/Longer Implementation: Increase focus on standardizing and measuring environmental variables in patients and model systems.

Day 1 | July 11, 2024

Welcoming Remarks

Amy S. Kelley, MD, MSHS, National Institute on Aging (NIA) and Michael Bennani, Ph.D., NIA

Dr. Kelley welcomed the in-person and online workshop participants. NIA leads the federal government in conducting and supporting research on health, well-being, and conditions associated with aging, including Alzheimer’s disease and Alzheimer’s disease and related dementias (AD/ADRD). To support these efforts, NIA is interested in data interoperability, which will provide valuable data infrastructure, support data accessibility, and accelerate aging and AD/ADRD research. The objectives of this workshop are to discuss the aging and AD/ADRD data ecosystem and consider what is necessary to establish an equitable and effective omics data ecosystem that will accelerate research discoveries. Dr. Kelley noted the breadth of perspectives and expertise represented in the workshop agenda. The input from this workshop will be integral to informing future strategic efforts in aging and AD/ADRD omics research. Dr. Bennani introduced the workshop co-chairs, Dr. Ohno-Machado and Dr. Bennett.

Workshop Background

Lucila Ohno-Machado, MD, Ph.D., M.B.A,. Yale University and David Bennett, M.D., Rush University

Dr. Ohno-Machado shared the key workshop themes covered in each session: (1) enhancing data interoperability, or the integration and accessibility of omics data; (2) focusing on diversity, equity, and inclusivity to address disparities in aging and AD/ADRD research; (3) improving global data sharing to overcome challenges and leverage opportunities for international collaboration; (4) tackling computational challenges associated with research on distributed data; and (5) enriching interoperability between model systems and human data to bridge gaps for better diagnostics and treatment.
This workshop was a follow-up from NIA’s FAIR (findability, accessibility, interoperability, and reusability) data practices workshop in June 2023 and shared examples of the challenges and opportunities for data infrastructure, interoperability, representation, equity, and inclusivity. Dr. Ohno-Machado introduced the chairs of each of the sessions, who provided overviews of their session through user stories that highlighted challenges and opportunities.

User Stories for Session 2: Diversity, Equity, and Inclusivity in Omics Research
Sid O’Bryant, Ph.D., University of North Texas Health Science Center

Despite the universal recognition of the importance of diversity, equity, and inclusivity in omics research, much work remains to be done. For example, the field has implemented more diverse model systems in aging and AD/ADRD omics research, yet many gaps persist in the diversity of human participants.

Additionally, reviewer feedback on grants and manuscripts does not reflect overall sentiment by the field to address diversity in these studies. As the field looks to improve diversity, equity, and inclusivity in aging and AD/ADRD omics research, it must consider five things: (1) identifying and gathering the necessary data, (2) ensuring the appropriate knowledge and resources are available to pursue certain research questions, (3) publishing and presenting findings, (4) sharing these findings with the broader community, and (5) educating the scientific community on the importance of diversity in omics research.

User Stories for Session 3: Global Data Sharing—Current Challenges and Future Opportunities
Heidi Sofia, Ph.D., National Center for Biotechnology Information (NCBI)

Data sharing has a close relationship to human rights. Article 27 of the Universal Declaration of Human Rights, created by the United Nations, states that “everyone has the right…to share in scientific advancement and its benefits.” Article 27 is a foundational and ethical principle of the Global Alliances for Genomics and Health (GA4GH), an organization dedicated to international cooperation on developing standards for data sharing. GA4GH is committed to unlocking the power of genomic data to benefit human health by establishing technical platforms and policy frameworks for sharing data.

One of the impacts of precision medicine is bringing the shared benefit of science back to the people. Genome-wide association studies (GWAS) are powerful analyses used to identify the genetic contribution to disease. Polygenic risk scores (PRSs) combine weak GWAS signals into a predictive score for disease and are commonly used by researchers; however, the lack of diverse representation in GWAS studies makes these PRSs much less reliable for underrepresented groups. Similarly, artificial intelligence (AI) models do not work well if they lack representative training data. Inclusion of diverse populations in genomics studies is also important to account for admixture (i.e., global relatedness between populations).

Importantly, precision medicine requires a variety of global data types besides genomics; these data include gene expression, multiomics, clinical records, environmental exposures, social determinants of health (SDOH), and real-world data. One approach for observing and collecting real-world data is through natural experiments, which are observational studies in which investigators cannot control certain factors (e.g., study of identical twins in different environments). Ensuring effective and ethical global data collection and sharing requires several important considerations, including building trust with research participants, establishing shared global data standards, advancing privacy with sociotechnical solutions, and ensuring shared benefits between researchers and research participants.

User Stories for Session 4: Omics Research in a Distributed Environment, Computational Challenges, and Solutions
Brian O’Connor, Ph.D., Nimbus Informatics, LLC

Dr. O’Connor shared two user stories focused on technical interoperability standards. The first user story focused on challenges with finding, working, and computing with data on distributed, disparate systems. Although several established standards can help with this issue (e.g., OpenID Connect, GA4GH passports), challenges persist with identity and access across international platforms, as well as with exploring data catalogs and searching for data in a centralized way. Standards like GA4GH Data Connect and Fast Healthcare Interoperability Resources (FHIR) have helped researchers more easily search and exchange data in a standardized way, but most data are only searchable within their specific programmatic spaces. Researchers may also experience challenges with taking data to a workspace environment where they can perform analyses. The NIH Cloud-Based Platform Interoperability (NCPI) effort is looking to standardize how data are being handled between platforms, and the GA4GH Data Repository Service (DRS) Standard is supporting secure access of data across platforms. Despite these efforts, there is no universal standard for using data references across platforms. Finally, tools and standards (e.g., GA4GH) support algorithm interoperability; however, many different workspaces use different algorithm wrapper technologies that make it challenging to port the algorithm from one location to another.

The second user story focused on harmonization and analysis, or how a researcher can work with a multitude of distributed multimodal data representing similar concepts that are encoded differently. Harmonizing data from a variety of sources using common data elements (CDEs) requires a lot of time and effort. Also, researchers must ensure privacy preservation when harmonizing and linking data from multiple sources. Although there are multiple privacy preserving record linkage (PPRL) technologies across platforms, linking data while protecting privacy requires significant coordination. Finally, AI tools can help reduce the time and effort to harmonize and link datasets for analysis. A multitude of AI services and customized models are emerging that have the potential to redefine the time-intensive tasks of harmonization, but they are not part of standard practice.

User Stories for Session 5: Creating Interoperability Between Model System and Human Omics Data in Aging and AD/ADRD
Melissa Haendel, Ph.D., FACMI, University of North Carolina

Dr. Haendel noted that mechanistic hypotheses of AD onset involving environmental exposures
(e.g., pesticide exposure) are difficult to study in comparative model systems and human populations. Model systems need to recapitulate the phenotypic and omic changes associated with the development and progression of AD. Phenotypic measures in model systems need to recapitulate human disease and be represented using ontologies and terminologies. Various phenotyping measures in humans can ensure interoperability with model systems, including clinical instrument measures, CDEs, and wearable data.

Modifiers of human AD cases can be captured through patients’ environmental exposures
(e.g., residential or occupational history) and other SDOH; however, capturing similar environmental exposures in model systems is challenging, especially since there are spatial and temporal considerations. Therefore, different models can recapitulate different aspects of AD. Although some core model organisms of AD exist, several natural models of AD can lead to a better understanding of the early influences that influence AD development and progression.

Session 1: Introductions to Aging and AD/ADRD Data Platforms and Other Data Ecosystems

Co-Chairs: Michael Bennani, Ph.D., NIA, and Mette Peters, Ph.D., NIA

National Institute on Aging Genetics of Alzheimer’s Disease Data Storage Site (NIAGADS)
Li-San Wang, Ph.D., University of Pennsylvania

NIAGADS is the NIA-designated national data repository for aging and AD/ADRD genetic data and facilitates qualified access and sharing of these data. NIAGADS also serves as the data coordinating center for the Alzheimer’s Disease Sequencing Project (ADSP) and shares data generated by the ADSP with the research community through the NIAGADS Data Sharing Service . NIAGADS has 125 datasets from 86 international cohorts with more than 200,000 samples and 23 datatypes, including single nucleotide polymorphism (SNP) array, whole-genome sequencing (WGS), and whole exome sequencing (WES) data. The goal of ADSP is to use WGS to lead to new insights in Alzheimer’s biology, target discovery, and genetics-driven clinical trials. Along with NIAGADS, ADSP is composed of several initiatives, including the AI/Machine Learning (ML) Consortium, the Functional Genomics Consortium, and the Follow-Up Sequencing Study, which is looking to increase diversity of the cohorts in the dataset. ADSP’s Phenotype Harmonization Consortium (PHC) is harmonizing datatypes like autopsy, cognition, and imaging data from cohorts included in the ADSP dataset. The current ADSP data release (ADSP R4) includes WGS from 36,000 participants. The ADSP R4 dataset also has detailed quality controls, additional harmonized phenotypes, individual-level structured variant calls, and SNP array data from over 30,000 samples from the NIA-funded Alzheimer’s Disease Research Centers (ADRCs). Importantly, approximately 55% of the samples in the ADSP R4 dataset come from non-White participants. The upcoming ADSP R5 data release will include data from more than 60,000 participants with strong racial/ethnic diversity.

NIAGADS follows the NIH Data Management and Sharing (DMS) Policy and the NIH Genomic Data Sharing (GDS) Policy . The Institutional Review Board (IRB) at the cohort study’s institution must approve inclusion of the cohort’s data in NIAGADS based on the research use limitation information in the original informed consent. This allows data from older cohort studies to be included in NIAGADS while respecting the original concept of the study. As dictated by the GDS policy, the NIAGADS Data Access Committee reviews and approves data access requests submitted from both NIH-funded and non-NIH- funded investigators who are seeking access to NIAGADS data and datasets.

Several other resources on the NIAGADS website support aging and AD/ADRD researchers. In addition to the controlled-access data, NIAGADS provides access to 8,000 open-access data with primarily summary statistics and related content that can be downloaded directly without submitting a data access request. The NIAGADS Alzheimer’s Genomics Database is one of the many genomics resources available through NIAGADS and partner sites and is an interactive knowledgebase platform for AD genetics. This database has summary statistic datasets for gene variant reports, annotations, and allele frequencies and enables data sharing, discovery, and analysis. The Functional Genomics Repository is a harmonized collection of functional genomics that tracks annotations for AI/ML research from across more than 20 primary data sources. Finally, the Alzheimer’s Disease Variant Portal is a harmonized disease-centric collection of high- quality and suggestive AD-genetic association findings curated from the literature. All these resources and other aspects of NIAGADS have features that support NIA’s FAIR principles (e.g., digital object identifiers for shared datasets, outreach initiatives).

AD Knowledge Portal and Exceptional Longevity Translational Resources (ELITE) Portal
Anna Greenwood, Ph.D., Sage Bionetworks

The AD Knowledge Portal and the ELITE Portal are NIA-funded data repositories that support aging, AD/ADRD, and longevity translational research at NIA. Sage Bionetworks serves as the data coordination center for these portals, meaning it supports data and metadata ingress, organization, and governance on the AD Knowledge and the ELITE portals. The AD Knowledge Portal is a data repository for 15 programs, including Accelerating Medicines Partnership® Program for Alzheimer’s Disease (AMP-AD). The portal contains information about programmatic goals, methodological data, contributing researchers, experimental models, computational tools, and other resources. The AD Knowledge Portal also shares community contributed data from researchers external to these 15 programs. The AD Knowledge Portal has over 23,000 data files, including data from human subjects (e.g., brain, plasma, cerebrospinal fluid) and model systems (e.g., mice and Drosophila). The portal contains diverse data types, including a variety of omics data, behavior data, and imaging data. In the future, new data types like exposomic, SDOH, and biomarker data will be added to the AD Knowledge Portal. The ELITE Portal has omics and model system data from three longitudinal cohort studies: Longevity Consortium, Long Life Family Study, and the Integrative Longevity Omics study.

Both the AD Knowledge and ELITE portals have a standardized explorer menu and use metadata annotations that have been adapted from community standards. Also, users can search for data based on specific participant characteristics, including age, diagnosis, race, and ethnicity, on the ELITE Portal and soon on the AD Knowledge Portal. Like NIAGADS, these portals use streamlined data access requests that are based on NIH GDS and DMS policy. Within the AD Knowledge Portal, users can access all the controlled data through a single data use certificate where all data is available for General Research Use. For the ELITE Portal, access to certain data types requires IRB review and approval. To improve interoperability, Sage Bionetworks uses some of the community standards, connecting several different compute platforms (e.g., CAVATICA, Terra, the AD Workbench), and is working with other NIH data repositories to improve data interoperability.

Alzheimer’s Disease Data Initiative (ADDI)
Matt Clement, Ph.D., ADDI and Mukta Phatak, Ph.D., Gates Ventures

ADDI is accelerating new discoveries by informing the development of diagnostics and treatments for AD/ADRD and enhancing the data ecosystem for a global community of researchers. ADDI supports the AD Workbench , which allows researchers to search datasets from a variety of sources and submit requests to access and analyze these datasets in cloud-based, collaborative workspaces. The AD Workbench has NIH-supported data sharing resources and tools to enable federate access to datasets. The AD Workbench’s Federate Data Sharing Appliance provides a foundation for secure data sharing and empowers researchers by unlocking access to previously unreachable datasets. This Linux-based system facilitates interoperability across different platforms without the data leaving the data owner’s environment. While users can search and view summary-level data about the available datasets, they must submit a request to the data contributor to access individual-level data. Once approved, users can access these data in the collaborative workspaces, which have standard analysis tools (e.g., Jupyter Notebooks) and unique software tools like a biomarker analysis tool from Roche. Researchers can search and submit access requests to over 75 datasets on the AD Workbench, including long-read DNA sequencing, RNA sequencing, and proteomics data. Several partners using the AD Workbench are analyzing data within the collaborative workspaces and plan to share these workspaces with other users.

NCBI database of Genotypes and Phenotypes (dbGaP)
Kim Pruitt, Ph.D., NCBI

NCBI’s dbGaP is an archive of DNA genotypes and phenotypes from NIH-funded studies that represent more than 4 million study participants from over 2,700 studies. This includes over 12,000 phenotype files and over 2 million sequencing files. The dbGaP also cross-links with other data repositories like ClinicalTrials.gov, PubMed, and ClinVar. Summary-level data is publicly available, but individual-level data can only be accessed through a controlled access process. The dbGaP is part of the NCPI program. As part of developing dbGaP, NCBI worked with the NIH Center for Information Technology to define the technical specifications for the NIH Researcher Auth Service (RAS) , which provided a centralized NIH solution for authorization and authentication of controlled access human phenotype and genotype data across the NIH federated landscape. As a result, the dbGaP was the first federated data repository to utilize and provide authorization to RAS. Along with NIH RAS, dbGaP uses GA4GH DRS and the HL7 FHIR technical standards. The dbGaP manages data access requests, datasets being added to the repository, the creation of metadata records (e.g., BioProject, BioSample) and the quality assurance and quality control (QA/QC) measures. Through partnerships with communities, dbGaP is committed to maintaining public trust with reliable, high-quality, FAIR data that empowers biomedical discovery accessible and user-friendly data and tools.

National Heart, Lung, and Blood Institute (NHLBI) BioData Catalyst (BDC)
Regina Bures, Ph.D., NHLBI

BDC is a community-driven ecosystem that is implementing data science solutions to democratize data and computational access to advance heart, lung, blood, and sleep (HLBS) science. BDC is a pillar of NHLBI’s overall strategy to support the research community by developing and integrating advanced infrastructure leading new tools and FAIR data. BDC has a wide range of users from citizen scientists to research scientists. Originally, BDC was built to focus on Trans-Omics for Precision Medicine (TOPMed), but it has since transformed to focus on multimodal data, including genomics, clinical, and imaging data from over 500,000 participants from 185 studies. BDC is a cloud-based, secure infrastructure that has several interoperable platforms. Within the BDC ecosystem, users can search available data, access workflows, and use the TOPMed Imputation Server. BDC has approximately 4 petabytes of HLBS data, supports team collaboration through its workspaces, and provides tools for analysis. NHLBI also allows researchers to bring their own data and tools to the BDC platform. NHLBI offers support, tutorials, and documentation to help users as well as $500 in cloud credits. BDC is leveraging interoperability in several ways. BDC is part of NCPI, which seeks to improve interoperability across NIH platforms, and the Advanced Research Projects Agency for Health (ARPA-H) Biomedical Data Fabric (BDF) Toolbox, which is an opportunity to leverage interoperability across federal agencies. BDC was recently named a GA4GH driver project to support its commitment to interoperability across national and international data sources.

National Human Genome Research Institute (NHGRI) Analysis, Visualization and Informatics Lab-space (AnVIL)
Chris Wellington, BSc, NHGRI

AnVIL is NHGRI’s cloud-based platform for secure data storage, analysis, and sharing that has more than 5 petabytes of data, specifically short-read WGS data from several different consortia. AnVIL seeks to address challenges from user journeys, including harmonizing of phenotypic data and molecular data, supporting data governance, adding value to existing datasets, supporting technical interoperability in a federated ecosystem, and encouraging analysis of diverse datasets. AnVIL is actively harmonizing data to its data model but remains flexible to accommodate the ever-evolving field of genomics.

AnVIL has two paths to data access. First, the submitter of the data is governed by their own data sharing agreements, and affirming compliance with their own policies is required. Second, there is NIH- adjudicated sharing through controlled access, which is controlled through dbGaP and adjudicated by NIH Data Access Committees. This data access path allows researchers to reshare data that they worked to harmonize within AnVIL. The compute infrastructure of AnVIL has standard analysis software but also allows users to bring their tools to the platform. AnVIL supports interoperability through NCPI and GA4GH and allows researchers to move data between different cloud providers and locations. AnVIL is continually working to improve its active user base through outreach and engagement with diverse audiences.

U.S. Department of Veterans Affairs’ (VA) Million Veteran Program (MVP)
Mark Logue, Ph.D., VA

MVP is a national research program funded by the VA Office of Research and Development. Its goal was to build one of the largest medical databases with one million veteran participants, and that goal was reached within the last year. The primary sources of phenotypic data for MVP participants include VA electronic health records (EHRs), surveys, and some data imported from external databases (e.g., National Death Index). The omics data in MVP includes genome-wide genotype data on 650,000 participants imputed to the African Genome Resources and TOPMed panels. MVP is generating WGS for approximately 150,000 participants, though researchers have access to only around 10,000 WGS and genome-wide DNA methylation data for 45,000 participants.

There are many data restrictions for MVP data; MVP is only accessible to VA-funded investigators with projects approved for MVP access. Phenotype development is handled in a separate system within the VA Informatics and Computing Infrastructure (VINCI) known as MVP-VINCI. For genetic analyses, phenotypes (e.g., case/control definitions) are transferred from MVP-VINCI to a separate MVP computing platform and assigned new identification codes. The VA’s central IRB requires that all identifiers be removed before MVP investigators can access the data to restrict recontacting participants for additional data collection. The VA Data Core must approve any data that is imported into MVP-VINCI from other sources. VA and MVP are working toward making MVP data access less restrictive through MVP Data Commons, which will allow non-VA investigators to analyze curated MVP phenotype and genotype data and receive summary output without directly accessing individual-level data files. The development of MVP Data Commons is ongoing and anticipated to be launched in 2025. Open access summary results from the MVP are available through dbGAP .

Very few autopsy- or biomarker-confirmed cases of AD are included in the VA EHRs. Often, EHRs have nonspecific dementia codes, VA patients receive care outside of the VA, or prescription of AD medication is nonspecific. Also, high rates of other dementias (e.g., vascular dementia) in the EHRs makes it difficult to determine the primary etiology. The MVP Cognitive Decline and Dementia During Aging Working Group is addressing these challenges with the MVP dataset. Early projects have focused on identifying AD, ADRD, and dementia cases within MVP using EHR and APOE genotype data. Of note, MVP has one of the largest genomic databases for people of African ancestry. The working group just launched several pilot projects to curate and harmonize data to the NIA ADSP in collaboration with the ADSP PHC.

Data for Indigenous Implementations, Interventions, and Innovations (D4I) Tribal Data Repository and the Rapid Acceleration of Diagnostics (RADx) Initiative
Matt Anderson, Ph.D., University of Wisconsin-Madison

RADx was established in response to the COVID-19 pandemic, and various projects within RADx, such as the RADx-Underserved Populations (RADx-UP) project , collected a large amount of data on a diverse group of participants. Eleven RADx-UP projects focused on American Indian/Alaska Native (AI/AN) populations collected a variety of data types that were stored in the RADx Data Hub. Through the efforts of Stanford University, Duke University, other scientific experts, and Tribal Nations, these AI/AN data from RADx-UP were transferred to the D4I Tribal Data Repository . Along with the FAIR principles, the D4I Tribal Data Repository also follows the CARE principles: collective benefit, authority to control, responsibility, and ethics. To support the CARE principle, the D4I Tribal Data Repository project goals are focused on supporting the administrative, governance, and Tribal consultation logistics; building the data infrastructure; and educating AI/AN scholars, Tribal government entities, Tribal communities, and non- AI/AN scholars who are interested in using AI/AN data for their research. D4I Tribal Data Repository governance is centered on data and engagement. Data and engagement committees include representatives from every participating Tribal Nation who are involved in the decision-making process and advisory groups that have experts in AI/AN data sharing and engagement.

The D4I Tribal Data Repository is collecting and aggregating data from RADx-UP and working with Tribal Nations to develop a framework by which AI/AN data can be added to the repository and made available for researchers. This will require many discussions, but the data will either be accessible under a federated system under Tribal jurisdiction or under a common unified model at NIH. Importantly, data from the repository will not be downloadable. The repository will also have a portal with a tracker of proposed projects, a dynamic consent portal where Tribes can opt in or out of research in real time, differential access based on account type (e.g., Tribal leader, research participant, researcher), and support pages for users (e.g., frequently asked questions, video walkthroughs). Overall, the goal of the repository is not just providing access to data, but rather engaging participants who provided these data throughout the process to ensure that they benefit from the research conducted using their data.

Audience Questions and Answers with Session 1 Presenters
Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

The presenters discussed important steps to create interoperable aging and AD/ADRD resources. These include having a data catalog, ensuring findability, and establishing standards at baseline. Developing the various aspects of interoperability standards (e.g., data management, data sharing policies) at baseline should be done collaboratively and in parallel. Also, researchers who run cohort studies or contribute to data repositories need to communicate with each other to ensure data are generated in common ways (e.g., mapped to the same genome build). Researchers should also consider the scientific questions these data will enable, which will help guide various interoperability steps (e.g., harmonization).

Efforts are underway to harmonize data access request processes and policies across aging and AD/ADRD data repositories. For example, individual-level data from studies registered within the dbGaP system cannot be re-distributed. The dbGaP acts as an entry point for users to request access to data and receive approval, but not all NIH-funded studies are registered in dbGaP. Importantly, data sharing occurs between institutions and not individual investigators. Various data repositories could leverage the NIH RAS and GA4GH to standardize data request processes across the aging and AD/ADRD research space.

Although cloud computing should provide equal opportunity for researchers to access and analyze data, some characteristics of cloud computing create barriers. First, cloud computing is expensive. Some type of support (e.g., free credits, funding) should be implemented to promote equity and increase accessibility to cloud compute services. Second, it is not always cost-effective to use cloud computing for certain tasks, like processing raw data into a derived dataset. There should be special considerations of how resources can be used more efficiently in a cloud compute environment. Groups should also consider the balance between improving the usability of the cloud computing infrastructure and ensuring these data and resources reach a wide range of users.

The presenters provided additional details about the various platforms. The D4I Tribal Data Repository supports various training efforts to reach people who may want to be involved in research, including the Summer Internship for Indigenous Peoples in Genomics and the annual IndigiData Science Workshop, which trains the next generation of AI/AN data scientists. Once the D4I Tribal Data Repository can be shared, this workshop will use RADx and repository data as part of its training. There may be other opportunities to provide additional training to other groups like community members, undergraduate students, and elected Tribal officials. The AD Knowledge and ELITE Portals allow users to download their data, but many researchers choose to work with the data in these portals’ cloud computing environments for ease and security. These portals follow the interoperability standards set by the GA4GH Data Repository Service and are working to connect with various compute platforms like CAVATICA, Terra, and the AD Workbench to democratizes access. Although MVP only has genome-wide genotypes, WGS, and DNA methylation data, inclusion of proteomic data is under discussion. Currently, only VA-affiliated investigators or researchers that collaborate with VA-affiliated investigators have access to MVP. These VA-affiliated investigators must have a VA-funded and approved project to access MVP.

Some datasets include multiomic data collected from the same sample or cohort. The NIH’s Accelerating Medicines Partnership is focused on having WGS, proteomic, metabolomic, transcriptomic, bulk sequencing, and single cell sequencing data for thousands of participants. ADSP, TOPMed, and AnVIL also have multiomic datasets on individuals.

The group discussed how to incentivize researchers to use a cloud workspace instead of downloading data and computing on premise. Institutions like NIA currently absorb the cost of users downloading data, but eventually institutions could require users to cover these costs. Trainings can educate researchers about the benefits of cloud-based research, which include collaboration and access to tools that may not be available at their institution. Workshops and pilot test programs can showcase how to use the cloud-based resources. Institutions could provide cloud compute credits to encourage researchers to use the cloud-based platforms.

Session 2: Diversity, Equity, and Inclusivity in Omics Research
Chair: Sid O’Bryant, Ph.D., University of North Texas Health Science Center

Many scientists are pursuing important research questions but often do not have the necessary data to answer these questions. Large data repositories provide these data to researchers without them having to collect the data themselves. Even though many research studies are not equipped to fully investigate health disparities, this type of research can be supported by ensuring studies are more inclusive of underrepresented populations. Many studies need to improve the collection of phenotypic data in diverse, underrepresented populations (e.g., racial and ethnic minority groups, sexual and gender minority groups), especially aspects of a person’s life like residential or job history. A major opportunity for the field is having a single resource that details what data are available from which repository or institution, which can help accelerate aging and AD/ADRD omics research.

Combining and harmonizing data from across many repositories remains a challenge. Collaboration between researchers with specific expertise is necessary to help researchers understand how to use and integrate certain datatypes into research projects. More team science-based approaches can help with answering these complex questions in an effective, appropriate way. Importantly, health disparities researchers should act as a resource to other researchers to help improve inclusivity and avoid biases or stigma. Like the proposed resource detailing available data, there could be a resource of researchers who are willing to share their expertise with others.

Even if they have novel findings, studies with small sample size or that are replicating findings in a different population may be difficult to publish. The field needs to recognize the importance of publishing work focused on factors across diverse populations. Similarly, the field needs to focus on providing results back to the people that participate in research and the communities they represent. Giving back to research participants is tremendously important and valuable, especially because the data is theirs, not the researchers. Finally, there are opportunities to educate the scientific community on the importance of diversity in omics research and datatypes like SDOH and race/ethnicity. Educating the scientific community on the importance of diversity can encourage the next generation of researchers to pursue this type of research. Importantly, the field should hold itself accountable to ensure that the diversity of the United States is reflected in their studies.

Genomics Through the Lens of Lived Experiences and Environmental Influences Affecting Communities with Varied Ancestral Backgrounds
Robert Barber, Ph.D., University of North Texas Health Science Center

Genetic ancestry is based on averages across the entire genome, but admixture within ancestral groups indicates that race and ethnicity are artificial rather than biological constructs. Consequently, if race and ethnicity are artificial, then differential risk and disease etiology between populations is also artificial. The focus of research should be on factors that drive disease risk like SDOH, which can be experienced differently across populations. These SDOH can have various impacts on biology, such as oxidative damage to DNA, changes in gene expression, or impacts on immune response, and can generate an individualized rate of gaining that impacts lifespan or risk for diseases like AD.

A person’s genetic predisposition for disease combined with lived experience and environmental exposures directly impact transcriptomic outcomes, proteomic outcomes, gene regulation, and ultimately phenotype. Research can focus on slowing the rate of AD progression by modifying environmental factors. Multiple cohorts and studies have demonstrated the correlation between DNA methylation and aging; however, the underlying mechanisms is not fully understood. What is known is that healthy lifestyle choices slow aging, and poor lifestyle choices accelerate aging. A recent study found that the rate of aging based on DNA methylation between non-Hispanic White, Black, and Hispanic differed; however, the use of different aging rate algorithms produces different outcomes for each group. This suggests that there is not enough data to accurately answer this question. Importantly, other studies in cardiovascular disease and cancer have shown that DNA methylation mediates SDOH and risk factors for disease. Though the more data types are needed (e.g., multiomics) to determine if this is the same for AD/ADRD, the field really needs more inclusive cohorts and better ways to harmonize the data.

Assessment of Redlining Practices that Impact Health Outcomes in Highly Vulnerable Neighborhoods, Data Analytics, Algorithmic Biases, and Research Integrity with Emphasis on Social Vulnerability and the Integration of Genomics with Non-Genomics Data to Shape Scientific Knowledge
Latrice Landry, Ph.D., University of Pennsylvania

Many influences from a person’s lived experience that contribute to the phenotype of diseases (e.g., environment, society, family, diet) are not considered or built into the clinical genomics research infrastructure. Models like Bronfenbrenner’s ecological systems theory account for all the environments (e.g., school family, neighborhood, community) a person directly or indirectly encounters in their lifetime that can impact their health. To add to this complexity, people have different experiences within these systems. Researchers need to consider how to capture and integrate these complexities into the clinical genomic research. Biomedical and clinical research can study these ecological systems using the translational science principles developed by the National Center for Advancing Translational Sciences (NCATS). Another tool is the translational research framework in genomics, which considers moving research questions through a pipeline that goes from basic research, translational research, translation in patients (i.e., clinical trials), translation into practice (e.g., human services research), and translation into community (e.g., policy, public health). This framework has recently been revamped to mirror the ecological systems model and includes the concepts of diversity, equity, and inclusion and the importance of ensuring these results can be easily disseminated and can be useful across populations.

Even though these tools are available, biases are already embedded in systems and approaches. This begins with problem selection, or what diseases are being studied or populations that are being included in these studies. The biases from problem selection impact data collection (i.e., data missingness), outcome definition, algorithm development, and post-development considerations. Therefore, these biases need to be considered across the translational research framework. Overall, there needs to be transdisciplinary approaches to including lived experience in genomics research. Redlining data and social vulnerability index (SVI) data are useful datatypes for this type of research. Redlining was a discriminatory practice of classifying neighborhoods as “desirable” or “undesirable” based on the populations living in those neighborhoods. Researchers have been using the redlining maps and integrating it with EHRs to build connections between health and environmental influences. The SVI uses variables from census data (e.g., poverty, age, disability, race and ethnicity, housing type) to identify communities that may need support before, during, or after disasters. Redlining and SVI data can be used to build hierarchical models that include individual- and community-level data, all of which help provide a more wholistic view of the influences on health and disease.

Community Values and Efforts to Communicate Research Findings from Omics Studies to Research Participant Communities
Barbara Harrison, M.S., CGC, Howard University

Sharing the results of genomic and other omics studies with research participants must be done in a way that mitigates harm. In the past, researchers have not always prioritized research participants’ well- being, which has resulted in many negative outcomes. There are many examples of scientific research harming individuals, such as using participant samples to study certain conditions without their consent (e.g., Arizona State University and the Havasupai Tribe) or not informing participants about research progress (e.g., sickle cell disease studies). These experiences have caused certain populations, particularly racial and ethnic minorities, to distrust or have no interest in participating in research studies.

Community-based participatory research (CBPR) engages communities and community leaders as collaborators with research teams. This involvement should begin at the start of the research process, before protocol development and funding awards. Key CBPR principles include dissemination, cultural competency, transparency, and capacity. Involvement from community members is also vital for communicating study results. Community leaders and members can help ensure that results are interpreted in a nonstigmatizing manner and use appropriate descriptive language. Researchers should also acknowledge the role of community partners in their publications by listing them as authors and work with these partners to develop plans for disseminating these results to the community (e.g., social media, radio, community events). There should also be efforts to provide information and resources to communities on how they can act on study results. Importantly, the relationships needed for CBPR take time to build and to maintain over the long term. The field must move away from the standard approach of research teams coming to communities, enrolling participants, collecting data, and leaving. These data belong to participants and not to researchers, so the field must dedicate itself to ensuring that these results get back to participants in an understandable and actionable way.

User Experience Challenges Facing Diverse Omics Data End-Users and Its Impact on Research While Addressing Viable Solutions that May Increase Data Access
Krystal Tsosie, Ph.D., M.P.H., M.A., Arizona State University

Although Indigenous people (e.g., AI/AN groups) are admixed populations, they are often categorized as genetically homogeneous. Reference genomes with Indigenous data include only a small number Indigenous people from diverse geographic locations (e.g., North America, Central America, South America). Meaningful genetic variant calls in small, admixed populations are lost to standard QA/QC processes. Importantly, focusing on biologically pure or least admixed individuals ignores the lived experience for many Indigenous people. Even though the common sentiment is that diverse omic data are needed to advance genomic health equity, Indigenous people still represent less than 1% of research participants. To increase Indigenous people’s engagement with research, the field must overcome mistrust, prioritize benefit sharing, and empower data-decision equity. However, increasing diversity in omics datasets is not going to solve health equity problems. Therefore, the messaging that Indigenous communities will not benefit from precision medicine if they do not participate is coercive and misleading.

Indigenous and AI/AN genomic data sovereignty is the right of these Indigenous nations and Indigenous people to exercise autonomy to protect their interests related to genomic data. This right is intrinsically defined by the Indigenous community and not the government. Despite Indigenous genomic data sovereignty, researchers and institutions find loopholes (e.g., resequencing Indigenous samples because their collection predated Tribal IRBs). To avoid these loopholes, the field needs to rethink informed consent specifically for Indigenous people. Standard informed consent language centers on the risks and benefits on the individual, which is culturally inconsistent with the ethics that the Indigenous community delegates as a collective. Also, there is a false notion that removing personal identifiable information protects the data. Data can still be reidentified, which is especially dangerous for smaller communities like Indigenous people. Therefore, the current standard for individual informed consent for genomic studies is inadequate for Indigenous communities; consent should adequately disclose the group risk. Research teams should also reconsider the guiding principles for data governance and data access. While the FAIR principles are focused on the researchers’ access to the data, research teams should consider the CARE principles that focus on community benefit and collective responsibility.

The Native BioData Consortium is the only existing Indigenous biological data repository. The consortium utilizes Tribal IRBs, receives input from the Tribal community members, and focuses on research questions that Tribal members are interested in pursuing. Some digital tools developed by Native BioData that facilitate Indigenous genomic data sovereignty include blockchain, federated systems, metadata labeling, and dynamic consent portals that allow participants to be key decisionmakers. Overall, the field must account for several important facets when considering Indigenous people for research. First, recruitment is not the same as engagement. Second, the field should consider implementing responsible informed consent, data access, and data sharing principles that respect Indigenous people and acknowledge Indigenous data sovereignty. Finally, the field should ensure that innovation is aligned with equity. Without progress in these three equity pathways, the field will default to the status quo.

Audience Questions and Answers with Session 2 Presenters
Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

Participants and presenters discussed data governance for Indigenous populations that do not have Tribal councils (e.g., Indigenous populations in Central or South American countries). Researchers engaged with Indigenous populations should always consider lived experience and social climate of these countries; for example, structural racism may have an impact on resources like health care for these Indigenous people. When a government does not acknowledge an Indigenous population or their rights, researchers need to consider how to adequately receive community-level consent and agreement on data governance. Research teams may act as data stewards on behalf of the Indigenous community and must consider how to protect their data throughout the data lifecycle. One consideration is for researchers to agree not to release Indigenous data if it will be used for health research, since it can lead to stigmatization.

There may be instances where genetic analyses identify Indigenous ancestry in people who did not know previously or where people who self-identify as Hispanic or Latino discover they have Indigenous ancestry. NIH could facilitate conversations between researchers and Indigenous communities about how to appropriately handle these situations with participants before publishing or releasing data into a repository.

Overall, it is important to have conversations with communities before they contribute their data to research. They need to clearly understand how their data will be used, where it could be going (e.g., repository), and potential implications of research studies that use these data (e.g., technology, drug treatment, patentable or commercialized products). While there are some tensions between various populations and the scientific community, these tensions do not need to exist. Science does not have to be in opposition to community, but this may require doing science differently. This involves engaging communities early, even before a proposal is written and funds are allocated. Insights from the community can support a positive, safe experience for research participants and ensure that results are shared through effective channels to the community. The idea of giving back to the community through research results is a major area of opportunity for the aging and AD/ADRD research field.

A workshop participant from the United Kingdom shared their experience with participant involvement and engagement in research. Patients and members of the public who are involved in the research process are co-authors on papers, serve as members of data access committees, and help formulate research questions. This is becoming standard practice at research institutions as well as charities that support various scientific topics. Along with promoting participant involvement and engagement, researchers should promote transparency. Research institutions in the United Kingdom are required by the government to publish data use registers to ensure ethical and responsible use of participant data. These registers include information about who accessed the data, why it was accessed, what publications used these data, and any secondary or tertiary research coming from these publications. These registers are an important part of the research infrastructure that supports transparency. The European Union has the General Data Protection Regulation (GDPR), which is a privacy framework that support ethical and responsible handling of participant data. The US should establish a centralized privacy framework, rather than having researchers navigate a variety of different privacy regulations.

There are several ways to encourage researchers to increase representation in their research studies. One way is implementing milestones that researchers must establish when applying for funding and report on progress throughout the funding period, like NIH’s UH2/UH3 and K99/K00 awards. Additional funding is not provided if certain milestones are not reached. This could be a standard for all NIH-funded research to hold researchers and funders accountable for increasing representation in research. As researchers try to meet communities where they are (e.g., clinical trials at community health or federally qualified health centers), NIH could develop an infrastructure grant that supports research teams and subject matter experts who are working to engage with communities and increase diverse representation in studies. This grant mechanism should include a monitoring and evaluation plan.

Redlining is just one of several aspects of structural racism that drive health inequities, though it is as not as prevalent today. Although capturing other data types related to structural racism is difficult, researchers can use redlining maps, census maps, environmental data, and EHR data to build models to understand how environmental factors play a role in disease.

One of the biggest barriers to achieving diversity, equity, and inclusivity in aging and AD/ADRD omics research is the lack of awareness within the scientific community on the importance of increasing representation and how to effectively engage underrepresented communities. Another major barrier is research funding. Funding is needed to support these community engagement efforts and train researchers in qualitative research, ethical data standards and practices, and other components of research focused on diverse, underrepresented populations (e.g., racial and ethnic minority groups).
Researchers are often not aware that funders like NIH will support grants that include community engagement plans, and so they do not include these plans in their budget justification submissions. There could also be opportunities for data repositories to require researchers to complete trainings on topics like participant privacy and ethical data sharing.

Session 3: Global Data Sharing, Current Challenges, and Future Opportunities

Chair: Heidi Sofia, Ph.D., National Center for Biotechnology Information (NCBI)

Enhancing Precision Medicine in Latin America: Ancestry, Admixture, and Genomic Data Sharing
Iscia Lopes-Cendes, MD, Ph.D., University of Campinas, Brazil

The Brazilian Initiative of Precision Medicine (BIPMed) was developed in 2015 to research complex diseases in admixed Brazilian populations, develop a reference dataset to account for this admixture, improve imputation quality, and establish parameters to classify variants (e.g., benign, likely pathogenic, pathogenic). Overall, the mission of BIPMed is to support genomic research, implement precision medicine, and improve genetic testing in Brazil. BIPMed hosts and shares genomic data from the Brazilian population and has WES and SNP array reference databases. BIPMed is part of GA4GH’s Beacon project, the National Initiatives Forum, and LatinGen.

Studies using BIPMed data found that almost 15% of rare variants identified in this Brazilian reference population were not previously reported. A comparison of common and rare variants (as measured by minimum allele frequencies) in the BIPMed SNP array reference database with common and rare variants in other population reference databases found that rare variants in the BIPMed reference were common variants in other populations. An admixture plot shows distinct differences between the Brazilian population and admixed American populations, which include Mexican, Columbian, and Puerto Rican populations. Together, these results indicate that each Latino or Hispanic population is distinct and should not be treated as a single population. There are even genetic ancestral distinctions within the BIPMed Brazilian population based on the regional area of Brazil. But despite having similar global ancestry in the Brazilian population, analysis of individual Brazilian data shows a mosaic of haplotypes from different ancestries on each chromosome.

BIPMed has established several efforts to protect participant privacy. While the risks associated with privacy and data sharing cannot be eliminated, they can be minimized. All BIPMed participants sign a consent that explains risks and potential uses of their data, and BIPMed follows local laws and regulations for privacy protection, which include anonymization and data encryption. BIPMed is part of the Beacon project and has different levels of access. While there is publicly available summary-level data, researchers must apply and sign a data sharing agreement to access individual-level genomic data.

Federated Data Access and Computing in Africa
Takudzwa Nyasha Musarurwa, M.S., University of Cape Town, South Africa

eLwazi is the open data science platform for the Data Science for Health Discovery and Innovation in Africa (DS-I Africa) consortium projects. The goal of eLwazi is to ensure DS-I Africa partners can find and access data, select tools and workflows, collaborate with each other, and run analyses on different computing environments. Currently, DS-I Africa researchers use and analyze data on the eLwazi platform in a nonstandard, noncollaborative way. DS-I Africa is updating the eLwazi platform to provide access to training materials, public data portals, and workspaces where users can save data from the eLwazi catalog. The updated platform will also allow researchers to select datasets, tools, and compute services. The platform leverages Terra and Gen3 implementations of the Data Biosphere concept and incorporates GA4GH standards. DS-I Africa’s research partners include institutions from Africa, the United Kingdom, and the US. DS-I Africa projects use a variety of data types (e.g., clinical, phenotypic, genomic) that require a standardized way of sharing. DS-I Africa institutions are from different countries subject to different data sharing and access policies. Therefore, the eLwazi platform needs to support standardized data sharing and collaboration. In August 2023, DS-I Africa hosted a GA4GH starter kit workshop that focused on introductory implementation to DRS , Data Connect , and Workflow Execution Service standards. As a result of the workshop, DS-I Africa began implementing starter kits on the eLwazi platform through pilot projects. These pilot projects are proofs of concept for data sharing and compute across different facilities using GA4GH standards. The first pilot project tested DRS servers at four African institutions and involved users querying the Data Connect servers, selecting the DRS Compressed Reference-oriented Alignment Map (CRAM) files based on the query, and submitting IDs for that data to a workflow execution service endpoint. The result was a multi-QC analysis that was generated using the CRAM files. This pilot testing allowed eLwazi to test the workflow using starter kits and eventually implement the full software into the workflow. The DS-I Africa team is testing more complex outputs that involve pulling from multiple datasets, like reference panels and imputation services. Overall, this development process allows the team to explore other GA4GH data security standards that can be implemented into eLwazi projects to protect data that is being processed and analyzed in different computing environments.

Privacy-Preserving Federated Analytics in Swiss Hospitals and Beyond
Jean-Pierre Hubaux, Dr.Eng., École Polytechnique Fédérale de Lausanne (EPFL), Switzerland

University hospitals across Switzerland were interested in exploring and analyzing siloed medical data without moving it. The solution was a privacy-preserving federated AI tool. To support this process, privacy enhancing technologies (PETs) were leveraged to protect data used as inputs for the analysis, including secure multiparty computation and homomorphic encryption . Secure multiparty computation means that data contributors complete computations based on other contributors’ inputs without revealing their own inputs. This approach is not completely secure (e.g., two contributors could collude and share inputs), so it can be combined with another PET known as homomorphic encryptions, which enables computations directly on encrypted data without having to decrypt it. These PETs were included in the federated AI tool at each data provider site so that computations can be carried out with the appropriate protections.

The first test version of this tool, which is called Tune Insight , was deployed at three Swiss hospitals and was later expanded to include a pharmaceutical company to establish a federated confidential computing system. A 2021 publication showed that Tune Insight can accurately reproduce survival analysis and GWAS results without moving the data or revealing inputs. The PETs used in Tune Insight have been reviewed by a multidisciplinary group of researchers along with lawyers to ensure they comply with GDPR. Along with hospitals, Tune Insight is being used by insurance companies, cybersecurity groups, and financial service groups. Overall, Tune Insight demonstrates the possibility of doing analytics on siloed data. This tool is being used for oncology and dermatology medical datasets and emergency pediatric data within the Swiss university hospital system as well as other countries outside of Switzerland.

Encouraging Open Data Provision: Challenges and Opportunities
Matteo Tranchero, M.S., University of Pennsylvania

Data collection and data sharing, also known as open data provision, are distinct and have different incentives. Researchers have strong incentives to collect data and publish research findings. Conversely, sharing data has no explicit incentives in terms of publications, career progression, tenure decision, or grant funding. Researchers are often hesitant to share data due to fears of being scooped, being proven wrong, or having data exploited. Research often focuses on data that are readily available. This translates into too much research on nonrepresentative populations (e.g., White, educated), and certain genes or diseases with fewer data available are neglected.

Based on this, open data provision is a market failure, which is an economic term meaning that the market forces caused under provision of goods relative to what would be socially optimal. For science, public data sharing clearly has more value for society than for scientists. The two solutions to fix this market failure are through public intervention (i.e., directly funding data collection and sharing) or regulations that force scientists to share data as a condition for publication or obtaining grants. The drawbacks of public intervention are disagreements around prioritizing which data to collect and share. Similarly, forcing scientists to share data may have unintended consequences, such as researchers avoiding publications or grant applications that have data sharing requirements. An important motivator in science is receiving credit for ideas through publications, but credit is really bestowed by the community of scientific peers who cite the publication. Bibliometrics is the quantitative count of citations of a given article, which signifies the support for a specific discovery by the scientific community. The issue is that citations are inherently social constructs and are often strategic.
Researchers tend to cite their advisors, collaborators, and editors and do not cite data sources. Although there are efforts to change the norms around citations and teach people how to cite public data, another solution is needed to capture credit for scientific ideas and innovation.

Logically, any scientific publication is composed of specific scientific entities, like genes or populations. Thus, data are codified measurements of these scientific entities; DNA is the sequence of a gene, and populations have specific characteristics. Therefore, the value of a scientific dataset can be based on the quantity and quality of follow-up research of a publication’s scientific entities, which is known as entitymetrics. Capturing the number of studies using certain entities is a new type of credit model that can also be used to show the impact of open data provision. AI/ML approaches can quantify entitymetrics by extracting information from follow-up studies that mention these entities. Importantly, entitymetrics is a more objective measurement of data impact that does not place weight on citations, is not limited to publications, and can use any output related to entities covered by the data, such as patents or clinical trials. Overall, AI/ML models that capture entitymetrics can be a more effective way to motivate public sharing of scientific data.

Audience Questions and Answers with Session 3 Presenters
Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA
The eLwazi data ecosystem is enabled by data passports for access to regulated data. This is a similar practice for other NIH programs, like the All of Us Research Program. This data passport access approach could be applied to other US data ecosystems, it just requires the proper infrastructure and expertise to establish. The eLwazi platform is a success story and clear proof of concept for the GA4GH starter kit. Other groups or institutions can replicate this work by reviewing the GA4GH standards or organizing a similar workshop. eLawzi investigators are also available to support and share their experiences. Groups can also review the eLwazi pilot project to establish GA4GH standard implementations on Github .

Within the credit allocation system established by entitymetrics, researchers could artificially increase the value of an entity’s data by controlling its market; however, the field can establish parameters around what counts as a downstream output enabled by the entity’s data. Importantly, bibliometrics can also be artificially increased by researchers using specific citations to boost their credit.

Research has shown that certain demographic groups of published scientists (e.g., Chinese researchers) are not always cited when they should be, and their data and results are used without the proper acknowledgement. The use of entitymetrics can identify publications that use certain entities even if they do not cite the original study. Entitymetrics could also be used to negotiate fair benefit for groups that donate data. Policymakers and users of the data can use entitymetrics to avoid misuse and capture the full potential of a given data source by tracking its uses, even if the data source is not cited. It can potentially provide a way to quantify the benefits of a given dataset.

The impact of admixture in the reference genome is not fully understood, because there is not much long read sequencing data for admixed genomes. Much genomic heterogeneity in populations has yet to be discovered because it has not been studied. Importantly, no single person will represent the variability in this admixed population, as is seen by differences in Brazilian individuals from different regions.

One of the main challenges with technological infrastructures adopting PETs include lack of awareness or understanding about PETs. There are also cultural gaps between computer scientists and lawyers, who are both committed to protecting data but through very different approaches. A third challenge with implementing PETs is lack of funding. Many times, organizations do not see the immediate benefit of investing in PETs other than promoting a reputation for data safety to the public.

Day 2 | July 12, 2024

Summary of Day 1 and Overview of Day 2

David Bennett, M.D., Rush University

Dr. Bennet, the co-chair of the workshop, provided a summary of day 1 presentations and discussions. He also shared a preview of the day 2 session and introduced the session chairs.

Session 4: Omics Research in a Distributed Environment, Computational Challenges and Solutions

Chair: Brian O’Connor, Ph.D., Nimbus Informatics, LLC

Dr. O’Connor revisited the user stories for Session 4 that were presented at the beginning of the workshop and highlighted how each presentation addressed these user stories.

The NIH Cloud Platform Interoperability (NCPI) Program
Valentina Di Francesco, M.S., NHGRI

NCPI is a partnership between cloud-based platforms across various Institutes and Centers (ICs) at NIH that seeks to develop and implement standards to enable interoperability and facilitate a NIH federated data ecosystem. NCPI is testing technical solutions and tools in production systems and leading conversations at NIH about policies and guidelines needed for implementing an interoperable federated data ecosystem. There are five partner systems within NCPI, including AnVIL, BDC, dbGaP, the National Cancer Institute (NCI) Cancer Research Data Commons (CRDC), and the Gabriella Miller Kids First Pediatric Research Program Data Resource Center (i.e., Kids First) from the NIH Common Fund. Together, these partner systems host 20.19 petabytes of data from 1.2 million participants. NCPI enables researchers to search for data from across these partner systems without having multiple logins and to access the results of these searches on a workspace of their choice. NCPI uses the previously established NIH RAS and DRS, which follow GA4GH standards, for data use ontology, genomics, federated search and discovery, and clinical and phenotypic data standardization. NCPI also uses the GA4GH tools Task Execution Service and Workflow Execution Service and is exploring other standards like FHIR. With support from the Office of Data Science Strategy, NCPI recently funded five interoperability projects that are focused on utilizing two or more participating NCPI systems.

NIH federated cloud-based data sharing has several noteworthy challenges. First, there are no standardized data element sharing agreements for open access, registered access, or controlled access. Second, despite the NIH Data Management and Sharing Policy that encourages researchers to share data, guidance is limited on how to share secondary data products, especially if these data are from controlled access datasets. While genomic summary results, imputation servers, and other forms for aggregate and imputed data are available, ICs have different policies about how controlled access data can be shared and under what conditions. Third, there are different views on data access for researchers developing and testing interoperability tools, including whether they need to submit data access requests. Also, there are disharmonious and incompatible authentication and authorization protocols between ICs. Finally, models of data stewardship, governance, and payment for cloud services differ between ICs. NCPI is leading conversations with ICs to address these challenges. Overall, NCPI is focused on addressing the technical and operational level of interoperability, which is foundational to interoperability.

ARPA-H Biomedical Data Fabric (BDF) Toolbox
Erika Kim, Ph.D., ARPA-H

ARPA-H BDF is a partnership program between NCI and ARPA-H to advance the next generation of tools to synthesize and speed the use of health research data. While BDF is initially focused on cancer, the goal is to generalize these tools for many different disease domains. The primary goals of BDF are to make biomedical research data easier to use, reduce effort for data integration, develop new BDF capabilities and tools, and build health data science models that can be applied across disciplines. Importantly, BDF provides a unified, consistent layer of data services that can work across many different systems and environments (e.g., data repository). The current challenges BDF aims to address are related to research data capture, EHR data extraction, data curation and harmonization, QA/QC, and user experiences with workspaces and data exploration tool. BDF is addressing and providing solutions to these challenges through five technical areas: (1) automated data collection, (2) AI-assisted curation, (3) intuitive exploration, (4), user testing, and (5) cross-domain generalization. The overall vision of BDF is to maximize the value and usability of clinical and experimental data for a variety of end users by enabling technologies from the first three technical areas while continually assessing and improving tools through technical area four.

The BDF tools will remove data sharing barriers, prepare data for analysis at scale, and provide continuous insight delivery through well-designed interfaces. The impact of BDF tools in the clinical setting include expediting clinical decision-making, improving diagnostics, and optimizing treatment plans. For research, BDF tools will advance the ability to integrate clinical data with experimental data that can lead to translation discoveries and build multi-institutional cohorts. For patients, BDF tools will improve their understanding of their health and treatment, identify relevant clinical trials, and help them to make more informed decisions. The work and tools of BDF will impact the community by encouraging community members to engage in research by helping them understand the impact of data sharing. To accomplish this, BDF is committed to having diverse use cases that will drive tool development. Many transition partners, including several NCPI platforms like BDC and AnVIL, will provide these diverse data use cases to BDF. These transition partners will also contribute data and tools to BDF as well as take prototypes of BDF tools and implement them on their platforms to test interoperability.

Compute Beyond the Shining Sea
Jack DiGiovanna, Ph.D., Velsera

Global collaborative research can help accelerate research on many topics, but there needs to be better international data sharing standards and policies. There are several examples of successful global collaborations. One example is Australia and the US sharing data from childhood cancer patients. Since high-risk childhood cancer cases are very rare, combining datasets across the globe can increase the power of studies. Zero Childhood Cancer in Australia and the Children’s Brain Tumor Network in the US were able to run RNA sequencing on the same platform by connecting the Australian and US compute platforms through AWS. While this pilot project was successful, it required international consortium agreements that took many months before this pilot could be started.

The work of NCPI has led to a significant evolution of interoperability standards, which has involved collaboration from many ICs and groups like Velsera and GA4GH. While there are many datasets that can easily be accessed and shared, there are data enclaves like MVP and the All of Us Research Program that are more restrictive. Storing data in an enclave helps to protect sensitive data or data from vulnerable populations. These enclaves have rich datasets that researchers want to access, but they present challenges for interoperability. A recent study comparing meta-analysis and pooled analysis approaches for analyzing genomic data from enclaves shows that the meta-analysis approach is a federate approach that brings compute to the data whereas pooled analysis brings the data to the compute like NCPI. Importantly, the meta-analysis approach is significantly more expensive and requires maintaining separate sets of codes to interpret the different workflow languages on each platform. All these challenges can be overcome by having the technology and policy sides of data sharing work together to promote interoperability that can lead to novel research questions and accelerate research.

The National Alzheimer’s Coordinating Center (NACC)’s Data Ecosystem, Revolutionizing Multimodal Data Integration, Interoperability, and Access to Advance Alzheimer’s Disease and Related Dementia Discovery and Translation
Sarah Biber, Ph.D., NACC

NACC serves as the data collaboration and communication hub and centralized data repository for NIA’s ADRC program and has become one of the largest and most comprehensive longitudinal, standardized, clinical, and neuropathological AD/ADRD datasets in the world. NACC has enabled major breakthroughs in the aging and AD/ADRD field and has been used to generate over 1,300 publications. Although data are standardized before they are submitted to NACC, NACC is actively working to establish pipelines and processes to integrate, harmonize, and share each of these data streams for the ADRC program. NACC is also working toward making a one-stop shop for ADRC participant data by building a cyberinfrastructure that expands interoperability and enables multimodal data integration, harmonization, and sharing.
Within this cyberinfrastructure, NACC launched a data platform that houses existing ADRC data and will integrate new data streams. Other efforts to standardize and streamline data include updating unique identifiers for each ADRC participant; assigning the PPRL as an additional identifier for each ADRC participant; establishing a new IRB to support collection, storage, and sharing of participant’s personal health information; and updating participant consents and research data use agreements to accommodate sharing new data streams.

NACC Data Front Door supports search, visualization, and access to all data modalities, including quick access files, real-time dashboards, secure enclave environment, and innovative core selection tools for building multimodal datasets. The quick access file feature improved the data request process and led to a 102% increase in data requests per year over the past two years. The real-time dashboards allow researchers to view the number of datatypes that have been provided by each ADRC, passed QC, undergone analysis, and are available to the research community. The secure enclave is undergoing pilot testing to allow researchers to combine research and real-world data, and the tools for building multimodal datasets are still under development. NACC is also looking to support computational analysis of these data on its data platform through a secure enclave environment for AI-driven discovery. NACC will leverage inputs from funded projects and research challenges to develop this enclave environment, which will help to establish a cloud compute credit model system to enable AI developments. Eventually, researchers can test and build off algorithms that were developed using NACC data. Overall, NACC is a unique and critical resource and has an important role to play in the evolving data ecosystem.

Leveraging AI to Harmonize Data at Scale
Shannon Ballard, Ph.D., Intramural Center for Alzheimer’s and Related Dementias (CARD) and Alan Long, M.S., CARD

DataTecnica and CARD continue to develop AI tools to support data harmonization from siloed research repositories. Two of these tools are Data Inventory and Validation Environment for Research (DIVER) and Generative CDE (GenCDE). DIVER allows researchers to browse various datasets using AI, and GenCDE allows researchers to browse thousands of generative common data elements (CDEs), which are common terms built from different data dictionaries using generative AI. DIVER and GenCDE systems first collect data elements in a variety of different formats from different silos and generate metadata for each input (e.g., expected ranges, likely abbreviations, or aliases). Next, these tools generate interoperability scores, which are quality reports of the data sets. These scores flag potential data coding issues and provide value information for the data scientists who are auditing these data and performing QA/QC. Federated analysis of these data can start once the data is matched. Importantly, the system learns over time and gets better at processing data from external silos.

DIVER helps researchers locate specific datasets from projects by managing and navigating large-scale data environments from multiple sources through a streamlined and searchable interface. This tool enables users to efficiently find and engage with the data they need for their specific research questions or projects. DIVER has a chat interface that allows researchers to directly ask questions about specific datasets, which humanizes the interaction and makes it accessible to all users regardless of their technical skills (e.g., using standard query language). The cataloging of data within DIVER is streamlined through a brief Google Form, which makes the process quick and more accommodating for users who may be less familiar with coding for data submissions. GenCDE complements DIVER’s functionalities by transforming sparse and messy data into structured usable formats using AI-driven algorithms and human oversight. This ensures accuracy and relevancy in the generated outputs. The capability of GenCDEs is invaluable for adapting to new research paradigms and data types, enabling the acceleration research outcomes, and contributing more effectively to the field. Data scientists review and audit harmonized datasets from DIVER and GenCDE achieve highest quality and reliability. Although this workflow is being used for scientific and health care data (e.g., Digitpath Research Sandbox), it can be deployed to areas like finance, climate change, and politics.

AI, Data Analytics, and the Integration of Environmental Data into Omic Datasets
Chirag Patel, Ph.D., Harvard Medical School

An early definition of the exposome was the time-dependent array of internal and external exposures humans encounter from birth to death. A recently developed definition for exposome is the integrated compilation of all physical, chemical, biological, and psychosocial influences that impact biology. By this definition, exposomics is a multidisciplinary study and a discovery-based analysis of environmental influences on health. The first main question for AD/ADRD exposomics research is understanding the amount of variation that can be attributed to the exposome in AD/ADRD. This involves computing the analog of heritability and determining the adequate power of studies to attribute certain factors to disease. The second question involves the factors of the exposome that are associated with AD/ADRD (e.g., environment, diet, lifestyle). Research answering this question could focus on which factors and how many are associated with AD/ADRD, how these factors work in aggregate (i.e., many factors with small effects or few factors with large effects on disease), and whether these factors are acute or chronic.

There are many data repositories, biobanks, and cohort studies with a variety of data types that can be leveraged to link genetic and environmental factors and understand their longitudinal influence on AD/ADRD. Along with getting access to these data, the challenge is choosing the right data for the analysis (e.g., studying location using ZIP code data or census tract-level data), which may impact the risk estimates, and phenotyping the outcomes (e.g., acute versus chronic disease). Another challenge is combining complex, heterogeneous modalities for multimodal exposomic research. These modalities can include high-resolution mass spectrometry (e.g., exogenous and endogenous compounds), targeted mass spectrometry (e.g., lead, cadmium), geospatial markers (e.g., ZIP code, air quality), self-report questionnaires (e.g., nutritional recall), untargeted mass spectrometry (e.g., mass-charge ratio), and sensor-based behaviors (e.g., accelerometers). Although there are many available knowledgebase resources available for exposomics research (e.g., Comparative Toxicogenomic Database), a limitation is understanding the evidence behind the relationships in those data. Multimodal AI and informatics methods can standardize the ways these heterogeneous exposomic and phenotypic data are collected and integrated at scale for these analyses. These methods are being tested and can lead to the development of atlases that make connections between exposomes and outcomes so that they can be queried and used for meta-analyses.

Audience Questions and Answers with Session 4 Presenters
Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

NCPI is continuing to address data interoperability issues within and across platforms, mainly by talking and collaborating with the IC groups running these platforms. Through these discussions, NCPI can either directly address technical challenges or raise issues with NIH leadership groups. NCPI can address technical issues by soliciting use cases, developing funding opportunities, or addressing these issues themselves. Although NCPI has not leveraged AI to address data interoperability issues, other groups at NIH are focusing on this issue. Outside of technical issues, NCPI will rely on NIH leadership groups to address policy issues (e.g., Scientific Data Council, Data Science Policy Council).

DIVER and GenCDE are currently in beta testing and not ready to be used by the broader researcher community, but CARD investigator Michael Nalls, Ph.D., can provide more details on the exact timing of releasing these resources to the research community.

Interoperability can unlock novel research questions and accelerate answers, which is an important return on investment for the field. Enhancing interoperability is a huge opportunity for NIH and NIA, and return on investment metrics can be used to understand how all aspects of interoperability such as diverse data contributions, compliance with DMS policy, and infrastructure building efforts translate into impactful research outcomes, as with the presentation on entitymetrics.

NACC has many unique aspects. First, the data are standardized at the ADRC program sites before they are incorporated onto the platform. Second, NACC has a large amount of longitudinal data, including a uniformed dataset of multidomain neurocognitive data, which is considered the gold standard for aging and AD/ADRD research and has been used by hundreds of international research studies. NACC is working to integrate other standardized data modalities within that uniformed dataset for each participant to make it more powerful. Also, NACC plans to use PPRL to allow cross-system analytics on sensitive data and bridge clinical and research data.

Early life exposures are an important consideration for aging-related conditions. Although most of the current exposomic data is related to current or recent exposures, lifetime exposures can be measured and analyzed, such as through lifetime residential history, biosamples collected early in life, and biomarkers that account for previous exposures (e.g., hemoglobin A1C). Researchers could also harmonize data from various cohorts at different stages of life. More efforts to collect early life exposures can be implemented to address these questions. Studies are underway to collect more detailed residential history and geocoded data.

The user testing technical area of the ARPA-H BDF toolbox will support challenges and other events like code-a-thons to test the newly developed BDF tools. These approaches will allow the community to provide feedback and vital input during the development of these tools.

Session 5: Creating Interoperability Between Model System and Human Omics Data in Aging in AD/ADRD

Chair: Melissa Haendel, Ph.D., FACMI, University of North Carolina

Most AD models focus on early-onset genetic disease and have had limited translational success. Model systems are also unable to recapitulate the environmental factors related to AD. The previous sessions demonstrated the advancing ability to study diverse populations, SDOH, and environmental factors, but challenges remain with studying these influences in AD model systems. The lack of interoperability between human studies and model systems in AD/ADRD has made it challenging to make correlations and therefore has limited the field’s ability to predict who will develop AD/ADRD and to create successful interventions to slow or prevent disease onset. One solution could be to have co-observational studies where observational cohorts are evaluated while AD/ADRD animal and cellular models are developed.

Translational Data Analysis for ADRD in the Model Organism Development and Evaluation for Late-onset Alzheimer’s Disease (MODEL-AD) Consortium
Gregory W. Carter, Ph.D., The Jackson Laboratory

Current AD mouse models emulate highly pathogenic mutations using human transgenes. The translational success of these mouse models has been limited, so MODEL-AD is working to develop new AD/ADRD mouse models that reflect the complex etiology of late-onset AD. MODEL-AD identifies variants of interest from large-scale GWAS based on multiple criteria (e.g., conserved between humans and mice, replicated in multiple studies) and integrates these variants into mice. Various phenotypes of these newly developed AD mouse models are characterized using omics (e.g., proteomics, transcriptomics), imaging, and neuropathology. Researchers can also test environmental exposures like diet in these AD mouse models to determine whether these factors drive additional phenotypes. Age- dependent changes can also be assessed in these AD mouse models, such as changes in gene expression. Characteristics of the older AD mice can be compared to data from endpoint AD studies (e.g., AMP-AD), but it is also possible to compare AD progression in mouse models and human patients using measures like proteomics and transcriptomics. MODEL-AD shares its raw data and protocols on the AD Knowledge Portal, summary target data on Agora, and interactive phenotype analyses of these mouse data on the Model AD Explorer.

Overall, accessible and reusable data from human models are essential for study design and analysis in mouse models. The goal of this work is to find AD phenotypes in the mouse models that emulate what is seen in human AD patients to support therapeutic development and testing. It is necessary to find the right model for the right treatment that can be administered at the right time in AD disease progression.

The Monarch Initiative, Leveraging Model Organisms to Characterize Phenotype to Disease Associations in AD
Monica Munoz-Torres, Ph.D., University of Colorado

The Monarch Initiative seeks to deliver knowledge about the relationship between genotype, phenotype, and disease using cross-species data. The Monarch Initiative Knowledge Graph is the aggregation of standards and tools that support semantic integration of these cross-species data from various sources. This unified data model brings together heterogeneous, multimodal, and disparate data in computable ways by unifying ontologies and leveraging semantics to support disease diagnosis, cross- species comparison, and treatment discovery. By making data more interoperable, the standards for data annotation and exchange help support data sharing and reuse by other projects and groups, which alleviates efforts toward data harmonization. Semantics acts as a universal convertor that allows information to be used for data models and other tools. Fuzzy matching, which is part of semantic searching, can be used across species to improve diagnostics and can reveal conserved mechanisms. For example, a study used the Monarch Knowledge Graph to predict factors associated with common female reproductive disorders (e.g., endometriosis) using harmonized survey data about internal and external environmental exposures and health conditions, biomedical ontology content, and supplementary nutrient and agricultural chemical data. The Monarch Initiative is beginning to leverage AI and large language models (LLMs) for curation, interpreting gene lists, parsing text to knowledge graphs, finding evidence in the literature, and mapping ontology terms.

Supporting LLMs with Structured Relationships to Integrate Knowledge of Models of Aging and AD
Harry Caufield, Ph.D., Lawrence Berkeley National Laboratory

Despite having incredibly large amounts of aging and AD/ADRD research data, challenges remain for comparing and reconciling observations across different model organisms, phenotypes, biomolecular metrics and assays, and environmental exposures. Traditionally, text mining has been used to compare observations. The first iterations of this text mining used rules and parsers, then moved to rule-based extractors, enrichments of terms and annotations, neural networks, and foundation language models. Each of these approaches is effective for different use cases, but very few work well for identifying relationships between concepts or events. LLMs offer statistical representation of text as data, but importantly are not based on fact. A recently developed LLM called the Structured Prompt Interrogation and Recursive Extraction of Semantics (SPIRES) is a knowledge extraction approach that is grounded in ontologies and other structure sources. SPIRES queries LLMS with schemas, which have concepts or entities to be extracted based on specific parameters. This LLM approach can be used to extract structured data from data models or from publications and develop a knowledge graph.
Dr. Caufield presented an example of this LLM approach using a schema for pulling information on AD/ADRD cases and pesticides in the literature. The schema includes information about the type of publication, model systems used, and concepts and relationships or associations mentioned in the text. The structured data output of this LLM quantifies the number of mentions of the criteria included in the schema, and these outputs can be integrated into relationships. Overall, LLMs are an improvement over previous approaches for several reasons. Researchers can modify schemas to fit different use cases without model tuning or retraining. Schemas can incorporate complicated relations constructed from concepts defined in the same or other schemas, LLM outputs contain document summaries and extracted terms that can be linked to ontologies that can easily be combined with other resources, including knowledge graphs.

Human iPSC-Based Experimental Systems for Capturing Genetic Diversity and Risk for AD/ADRDs
Tracy Young-Pearse, Ph.D., Harvard Medical School

iPSCs have become critical tools for studying neurological diseases that can complement existing animal model systems. iPSCs are pluripotent stem cells that are derived from human blood cells and can become any type of cells. Each AD/ADRD model is useful, but certain components of AD/ADRD can act differently depending on the model system, which is important to consider for interoperability. iPSCs are meant to model the human genetic aspect of AD/ADRD. Also, most model systems compare AD/ADRD with non-AD/ADRD; however, AD/ADRD is very heterogeneous in terms of age of onset, neuropathological burden, cognitive trajectory, and genetic risk factors. In fact, over 70 genetic loci are associated with late-onset AD, but only contribute to the heritability of AD. A large portion of AD risk heritability remains undefined due to many factors, including the lack of diverse populations in GWAS studies.

A diverse combination of genetic risk and resilience factors converge on interrelated but varied biological domains (e.g., lipid metabolism, synaptic vulnerability, microglial activation) to impact the risk and resilience of AD and cognitive and pathological traits associated with AD. Overall, all the genetic variants lead to the amyloid beta accumulation, tau accumulation, and synapse loss but are mediated by very different processes. iPSCs not only allow researchers to explore known and unknown genetic risks but provide a manipulatable system that can be used to actively measure different biological domains and test treatments to rescue these biological domains. The genetic and molecular characteristics of these patient-derived iPSCs can be associated with the neuropathology and cognitive trajectory of the donor. Also, the environment of iPSCs can be readily modified to define molecular consequences of exposure to neurotoxic agents and therapeutic interventions. There are also efforts to use iPSCs to improve understanding of the nongenetic influencers of disease, like bloodbrain barrier dysfunction. Overall, the heterogenous genetic and environmental risk factors associated with AD/ADRD disease pathogenesis can be modeled using iPSCs from diverse populations. Research using iPSCs can lead to the identification of drivers of disease and potential therapeutic targets.

Audience Questions and Answers with Session 5 Presenters
Michael Bennani, Ph.D., NIA and Mette Peters, Ph.D., NIA

The presenters discussed ways to address the “Aim 3 phenomenon,” which is the issue of modeling studies only being included in Aim 3 of the grant rather than being integrated throughout all proposed studies. Similarly, computational and interoperability resources are often limited to "Aim 3" of a project. Interoperability requires being proactive rather than reactive. For example, the usability of data generated from a study is limited if the data scientists are not included at the onset of a project. Researchers can leverage collaborations to integrate model systems more fully in their studies, but a multipronged approach may be needed to be mindful of the bandwidth of large consortia. Specifically, model system researchers should ensure that their model system can be applied to human research. Although the interoperability of molecular assays from models to humans is more straightforward, the field is beginning to do the same with behavioral and social assays. There should also be funding to support model system and human research collaborations; however, there are capacity limits for researchers to be involved in collaborations. Another consideration is using experts in semantic and knowledge engineering during the experimental design phase to ensure data produced can be interoperable and used for downstream applications. Grant applications should also require data management plans to ensure data produced from studies are usable and interoperable.

The Monarch Initiative and the Alliance for Genome Resources (AGR) are sister operations. AGR contributes data to the Monarch Initiative and utilizes many of its resources. The Monarch Initiative is a GA4GH driver project.

Many discussions at the workshop focused on data standards for patient data, but model systems need to be considered in these data standards as the field moves toward interoperability; however, these data standards will depend on the type of model system. Research involving iPSCs is already connected to human research, but research using animal models should consider assays or experiments that make the resulting data interoperable with human data. Behavioral and phenotypic data in humans are not easily translated into model systems, but this issue could be addressed through CDEs or semantics. A unified set of CDEs can support comparison and observation across species.

Workshop Summary

The chairs for sessions 2–5 compiled actionable items from their respective sessions, designating them by importance (higher or lower) and implementation timeline (quick or longer).

Actionable Items for Session 2: Diversity, Equity, and Inclusivity in Omics Research

  • Higher Importance/Quick Implementation: Establish a resource for investigators to identify experts and datasets to aid in the advancements of omics research among diverse communities.
  • Lower Importance/Quick Implementation: Establish a resource for investigators to identify resources to be able to conduct scientific analyses for the advancement of omic research among diverse communities.
  • Higher Importance/Longer Implementation: Establish a recruitment resource for existing NIA- funded programs to increase representation in clinical research.
  • Lower Importance/Longer Implementation: Establish a resource for modeling novel findings from diverse and representative communities in nonhuman systems.

Discussion

Working with diverse communities (e.g., racial and ethnic minority groups, sexual and gender minority groups) requires long-term engagement rather than short-term, one-off interactions. Sustained relationships with communities require community awareness, recruitment, and engagement. Although the aging and AD/ADRD field is adept at recruitment, more work is needed to support awareness and engagement. For example, awareness and engagement require educating the community about the research, which can be low cost and low effort. Also, empowering researchers to conduct research in diverse populations is critical and should not be overlooked. NIA and NIH could establish incentives for researchers to keep them accountable, such as milestones and loss of funding if these milestones are not met.

Actionable Items for Session 3: Global Data Sharing

  • Higher Importance/Quick Implementation: Drive broad adoption of GA4GH standards in technical platforms and policy frameworks.
  • Lower Importance/Quick Implementation: Contribute genomic data to the global pangenome
    reference.
  • Higher Importance/Longer Implementation: Implement privacy-enhancing technologies that are transparent to the user into data resources, and use entitymetrics to measure data value, improve credit models, and promote fair models of benefits.
  • Lower Importance/Longer Implementation: Link global data resources in an interoperable network using ARPA-H data services.

Discussion

Global data sharing ties into inclusive representation in research studies. Global data sharing should work toward building a social infrastructure that is focused on a return of benefit for the community. Also, efforts to encourage people to contribute their data can improve available pangenome reference data and the analyses that use them (e.g., variant calling). Many conversations are needed about how data from certain populations could and should be included in pangenome reference sets.

Actionable Items for Session 4: Omics Research in a Distributed Environment

Each of the following action items support technical interoperability standards:

  • Higher Importance/Quick Implementation: Implement RAS passports.
  • Lower Importance/Quick Implementation: Adopt DRS1.5.
  • Higher Importance/Longer Implementation: Establish federated compute systems.
  • Lower Importance/Longer Implementation: Create a data catalog.

Discussion

RAS passports provide authentication services for a variety of researchers and act as a user token for data access. Implementing RAS passports as a standard across the field must involve knowledge sharing from groups that use these passports to ensure they can be used easily and universally. Similarly, developing a data catalog can leverage the expertise of groups like the National Library of Medicine. RAS passports can still provide protections for participant data through data access levels.

Actionable Items for Session 5: Creating Interoperability Between Model System and Human Omics Data in Aging and AD/ADRD

  • Higher Importance/Quick Implementation: Design model system experiments to match real- world patient populations.
  • Lower Importance/Quick Implementation: Require AD programs to include collaboration with experts in standards and interoperability, especially phenotyping and biomarkers.
  • Higher Importance/Longer Implementation: Increase focus on standardizing environmental and social measures in patients and model systems and utilize natural, diverse, and richly longitudinal cohorts together with real-world and model system data.
  • Lower Importance/Longer Implementation: Increase focus on standardizing and measuring environmental variables in patients and model systems.

Discussion

There can be conflicts between developing data standards for well-established, well understood datatypes, a desire to support innovation, and establishing better data standards for datatypes. Semantics can support interoperability without fully standardizing a datatype. There are also inconsistencies with evaluating, validating, and maturing data standards. One aspect of the actionable item for data standards could be defining success metrics for interoperability and criteria for developing data standards. This can be done by understanding researcher and community needs related to data standards, such as data collection methods. Importantly, some people will not contribute data electronically or digitally because they are not technologically savvy or do not trust certain technologies. To meet participants where they are, data collection methods could be adapted, or researchers could educate and build trust with participants. Participants should also have more control over their data and how it will be shared, which requires educating participants about the potential ways their data will be shared and used for research. In general, people are more willing to share their data if they understand how their data will be used.

Final Thoughts

The session chairs reflected on the commonalities between their actionable items. These included interoperability standards, inclusivity of diverse populations, and centralized databases of resources and expertise. Additionally, building capabilities early and collaboratively, implementing evaluation and metrics, and promoting accessibility are vital for interoperability.

Closing Remarks

Jennie Larkin, Ph.D., NIA

Dr. Larkin noted that NIA plans to support conduct of a data landscape analysis, and a request for information was recently issued to solicit community input about how to conduct this analysis. Many groups within NIA are working to understand the scope of NIA-funded research.

Dr. Larkin thanked the workshop co-chairs, the session chairs, speakers, and NIA program staff who instrumental in facilitating the workshop, and the workshop participants for their contributions to this engaging and informative meeting. Despite each session covering very distinct topics, many common threads emerged that all led back aging and AD/ADRD research interoperability.

Acronyms

AD: Alzheimer's disease

ADDI: Alzheimer's Disease Data Initiative

ADRC: Alzheimer's Disease Research Center

ADRD: Alzheimer's disease and related dementias

ADSP: Alzheimer's Disease Sequencing Project

AGR: Alliance for Genome Resources

AI: Artificial intelligence

AI/AN: American Indian/Alaska Native

AnVIL: Analysis, Visualization and Informatics Lab-space

ARPA-H: Advanced Research Projects Agency for Health

BDC: BioData Catalyst

BIPMed: Brazilian Initiative of Precision Medicine

CARD: Center for Alzheimer's and Related Dementias

CARE: Collective benefit, authority to control, responsibility, ethics

CBPR: Community-based participatory research

CDE: Common data element

CRAM: Compressed Reference-oriented Alignment Map

CRDC: Cancer Research Data Commons

D4I: Data for Indigenous Implementations, Interventions, and Innovations

dbGaP: Database of Genotypes and Phenotypes

DIVER: Data file, Inventory, and Verification Environment for Research

DMS: Data Management and Sharing

DRS: Data Repository Service

DS-I Africa: Data Science for Health Discovery and Innovation in Africa

EHR: Electronic health records

ELITE: Exceptional Longevity Translational Resources

FAIR: Findability, accessibility, interoperability, and reusability

FHIR: Fast Healthcare Interoperability Resources

GA4GH: Global Alliances for Genomics and Health

GDPR: General Data Protection Regulation

GDS: Genomic Data Sharing

GenCDE: Generative Common Data Elements

GWAS: Genome-wide association study

HLBS: Heart, lung, blood, and sleep

IC: Institute and Center

IHS: Indian Health Service

iPSCS: Induced pluripotent stem cells

IRB: Institutional Review Board

LLM: Large language model

ML: Machine learning

MODEL-AD: Model Organism Development and Evaluation for Late-onset Alzheimer’s Disease

MVP: Million Veterans Program

NCATS: National Center for Advancing Translational Sciences

NCBI: National Center for Biotechnology Information

NCI: National Cancer Institute

NCPI: NIH Cloud-Based Platform Interoperability

NHGRI: National Human Genome Research Institute

NHLBI: National Heart, Lung, and Blood Institute

NIA: National Institute on Aging

NIAGADS: National Institute on Aging Genetic of Alzheimer's Disease Data Storage Site

PET: Privacy enhancing technology

PHC: Phenotype Harmonization Consortium

PPRL: Privacy preserving record linkage

PRS: Polygenic risk score

QA/QC: Quality assurance/quality control

RADx: Rapid Acceleration of Diagnostics

RADx-UP: Rapid Acceleration of Diagnostics Underserved Populations

RAS: Researcher Auth Service

SDOH: Social determinants of health

SNP: Single nucleotide polymorphism

SPIRES: Structured Prompt Interrogation and Recursive Extraction of Semantics

SVI: Social vulnerability index

TOPMed: Trans-Omics for Precision Medicine

VA: US Department of Veterans Affairs

VINCI: VA Informatics and Computing Infrastructure

Contact Information

Please contact Tiffany Rolle at tiffany.rolle@nih.gov for questions you may have about the workshop.