These foundations provide the ontology & knowledge graphs, quality, identity, governance, and orchestration controls needed to reliably build and operate an enterprise AI layer for discovery.
From Raw to Ready: Designing for Data Readiness at Scale
Data preparation isn't just a scripting problem; it's a systems design problem. This session steps back from individual tools and workflows to examine the architectural choices that determine how efficiently a team can move data from raw to ready. Attendees will see how decisions made at the pipeline and infrastructure level directly shape AI performance, reproducibility, and team velocity.
Managing Data at Scale: Hardware Considerations and Practical Approaches
This talk explores the challenges of managing large-scale data in today's evolving hardware landscape. We'll examine current trends in NVMe and memory availability, discuss how these factors shape data management decisions, and walk through a practical approach to handling data at scale. Attendees will leave with a clearer understanding of the tradeoffs involved and concrete strategies they can apply to their own environments.
Unlocking Public Data at Scale and Closing Remarks
An exploration of the challenges in sourcing, cleaning, and harmonizing large public datasets. This session will highlight common pitfalls, such as inconsistent formats, missing metadata, and data quality issues, and introduce an approach using OpenClaude to transform raw public data into structured, digestible datasets with standardized baseline metadata.
Data Representation & the Cost of Standardization
A deep dive into data modeling: the role of standard dictionaries, schemas, and the decisions that shape them. This talk examines the real trade-offs in harmonization: what you gain in interoperability and scalability, and what you lose along the way.
Trusted AI Starts with FAIR Data: Why Community, Semantics, and Collaboration Matter
AI in life sciences R&D is increasingly constrained not by models, but by the quality, interoperability, and governance of underlying data. This talk presents a perspective from the Pistoia Alliance, integrating insights from its AI and FAIR Communities of Experts. We show how data readiness, semantic interoperability, and governance are key to scalable AI, and why trusted AI needs FAIR data. We share perspectives from the Pharma General Ontology-Terminology project and the FAIR Personas - including the emerging “Agentic AI†persona - to illustrate the importance of shared data language for both human and AI. Cross-industry collaboration to establish common semantic assets becomes a strategic necessity in this context. Finally, a virtuous cycle may be emerging: FAIR data enables better AI, and AI increasingly contributes to FAIR data.