Half a million genomes. 1.5 billion variants. One breakthrough: we are all truly unique. Twenty years ago, the Human Genome Project took 13 years and $2.7B to sequence a single genome. Today? We can sequence a genome in less than 24 hours for under $1,000. Last week, UK Biobank released 490,640 whole genomes — the largest genetic dataset ever (Nature, 2025). What did we learn? • Each person carries 4–5 million variants • 76% appear in fewer than 10 people — your genome is almost entirely yours • 1 in 10 carries clinically actionable mutations where doctors can intervene today (e.g., BRCA1/2 for cancer, LDLR for heart disease) Why it matters: • Previous genetic tests captured ~6% of human variation. This dataset reveals 40× more • In non-coding regions — the biological switches controlling genes — researchers found 63 new disease associations • Adding 31,785 non-European genomes uncovered 82 disease links invisible in Eurocentric studies From genetics to health impact This transforms medicine today: • Prevention - Polygenic risk scores flag disease decades before symptoms • Diagnosis - Rare disease patients waiting years for answers finally find them • Treatment - Pharmacogenomics matches the right drug, right dose, to your genome The next frontier: genetics + everything else Genetics is the hardware. Health is the software running in real time. Your DNA is fixed, but biology is dynamic, shaped by: • Epigenetics: how environment and lifestyle switch genes on/off • Proteomics & metabolomics: molecular signals revealing your current health state • Digital biomarkers: continuous data from stress, sleep, glucose, heart rate • Stress biology & neuroendocrine signaling: how cortisol and brain-body responses reshape your health trajectory Layer these dynamic signals onto genetic foundations, power them with AI, and you create living health models, not just predicting disease, but understanding when, why, and how it manifests in YOU. The critical question? We've spent decades treating the "average patient" — who doesn't exist. Now we can better see each person as they truly are: biologically unique, dynamically changing, infinitely complex. The healthcare winners of the next decade won't just collect data: they'll integrate genetics, epigenetics, molecular and phenotypic tests, lifestyle, stress biology, and digital signals to deliver truly personalized, preventive care at scale. There is no "normal" genome, only 8 billion unique experiments in being human. And we just decoded the first half million. 👉 Which excites you more: knowing your genetic blueprint, or understanding how your daily choices rewrite it?
Data Analysis In Biology
Explore top LinkedIn content from expert professionals.
-
-
Can large language models be used in biotech? The short answer is yes. While LLMs are often associated with chatbots, their capabilities extend beyond that. In biotech, much of the data comes in the form of sequences – like nucleotides in DNA, or amino acids in proteins. Similar to sentences in natural language, these biological sequences have unique semantic meanings based on the arrangement of their components. When input data is fed into an LLM, a transformer converts these sequences into contextual vectors using its attention mechanism. This process allows the model to understand the context and relationships within the data, enabling it to predict subsequent elements. One such use case is prediction of neoantigens that enable targeting tumor cells in personalized cancer immunotherapies. Neoantigens are tumor-specific mutated peptides presented on the surface of tumor cells because they bind to human leukocyte antigen (HLA) molecules. LLMs can predict this binding affinity. This allows the development of personalized therapies that use the patient's own immune system to kill tumor cells without damaging healthy tissues.
-
🗨️ Just published in Nature Biotechnology: Our CellWhisperer AI enables chat-based analysis of single-cell sequencing data. You can talk to your cells & figure out the biology without writing any computer code. Paper link: https://lnkd.in/db3tvWeh. ⚙️ To get started, let’s find cells by typing into the CellWhisperer chat box. For example ‘Show me structural cells with immune functions’. CellWhisperer scores each transcriptome by how well it matches this textual query and colors by query match. 🔍 We investigate one of the identified cell clusters by selecting the cells & prompting CellWhisperer with ‘Describe these cells in detail’. This interactive workflow is enabled by seamless integration of the CellWhisperer AI chat box into a version of CELLxGENE Explorer. 🔬 You can easily query large transcriptome datasets for your favorite biological process using CellWhisperer. Just open Tabula Sapiens (https://lnkd.in/dMyUvqQH) or GEO (https://lnkd.in/dXnDGa9c) in CellWhisperer & type your query into the chat box – for example “infection”. 🆕 The CellWhisperer paper includes several new analyses beyond our 2024 bioRxiv preprint (https://lnkd.in/dKS6aGXi). For example, we used CellWhisperer for an AI-guided analysis of human organ development. 🚀 We also validated CellWhisperer’s chat-based analysis with conventional bioinformatics. CellWhisperer was >4x faster (and 10x cooler 😊). Our recommendation: Use CellWhisperer for dataset exploration – but statistics is still important to ensure rigor & reproducibility. 🪄 How does CellWhisperer work behind the scenes? We trained a multimodal AI that links transcriptomes and text, enabling free-text search and annotation of RNA profiles. And we connected this model to an LLM that we fine-tuned into a chat assistant for transcriptome data. 📚 We trained on >1 million bulk & pseudo-bulk transcriptomes with textual annotations that we AI-curated from GEO & CELLxGENE Census. Our training data is open source and useful for developing multimodal biomedical AI models and future bioinformatics research assistants. Ready to talk to cells? 🧬 Try the web app with public datasets: https://lnkd.in/diQcSite 🖥️ Analyze your own datasets: https://lnkd.in/dgdXyvbq 📖 Read the paper (open access): https://lnkd.in/dQgHkZdR 🧬 In summary, CellWhisperer introduces a chat-based way to explore scRNA-seq data. By enabling natural language analysis, it bridges biologists and bioinformaticians—paving the way for AI-driven bioinformatics assistants. 🤝 Huge thanks to the team! Moritz Schaefer & Peter Peneder with Daniel Malzl, Salvo Danilo Lombardo, Mihaela Peycheva, PhD, Jake Burton, Anna Hakobyan, Varun Sharma, Thomas Krausgruber, Celine Sin, Jörg Menche, Eleni Tomazou, Christoph Bock at CeMM, Medizinische Universität Wien, St. Anna Children's Cancer Research Institute (CCRI).
-
Stanford researchers just introduced an AI agent that autonomously discovers new biology from single-cell RNA-seq data. Despite the explosion of publicly available datasets, many remain under-analyzed due to technical barriers and limited human bandwidth. 𝗖𝗲𝗹𝗹𝗩𝗼𝘆𝗮𝗴𝗲𝗿 𝗶𝘀 𝗮 𝗟𝗟𝗠-𝗯𝗮𝘀𝗲𝗱 𝗮𝗴𝗲𝗻𝘁 𝘁𝗼 𝗯𝘂𝗶𝗹𝗱 𝗼𝗳𝗳 𝗽𝗮𝘀𝘁 𝗮𝗻𝗮𝗹𝘆𝘀𝗲𝘀 𝗮𝗻𝗱 𝗴𝗲𝗻𝗲𝗿𝗮𝘁𝗲 𝗻𝗼𝘃𝗲𝗹 𝗵𝘆𝗽𝗼𝘁𝗵𝗲𝘀𝗲𝘀 𝗳𝗿𝗼𝗺 𝘀𝗰𝗥𝗡𝗔-𝘀𝗲𝗾 𝗱𝗮𝘁𝗮: completely autonomously. 1. Outperformed GPT-4o and o3-mini by up to 20% in predicting real author-conducted analyses across 50 scRNA-seq papers in the new CellBench benchmark. 2. Uncovered a novel finding that CD8+ T cells in COVID-19 patients have significantly elevated pyroptosis scores (p = 0.001), a hypothesis not examined by the original study. 3. Discovered menstrual phase-specific receptor-ligand signaling between endometrial stromal fibroblasts and endothelial cells, validated through 40 tested gene pairs. 4. Revealed increased transcriptional noise with aging in specific brain cell types (microglia, oligodendrocytes), using data from a subventricular zone single-cell atlas. CellVoyager leverages a dual-loop architecture: an LLM planner (o3-mini or GPT-4o) for hypothesis and code generation and a vision-language model (VLM) for interpreting outputs. I thought this separation of planning vs. interpretation modules was cool because it reminded me of modular cognitive architectures in robotics (such as perception vs. planning agents) I also liked the self‐critique step immediately after generating each analysis plan which honestly I need to do more when I come up with scientific plans haha. Here's the awesome work: https://lnkd.in/eE9eqv6t Congrats to Samuel Alber, Bowen Chen, Eric Sun, Alina Isakova, Aaron Wilk, MD, PhD, and James Zou! I post my takes on the latest developments in health AI – 𝗰𝗼𝗻𝗻𝗲𝗰𝘁 𝘄𝗶𝘁𝗵 𝗺𝗲 𝘁𝗼 𝘀𝘁𝗮𝘆 𝘂𝗽𝗱𝗮𝘁𝗲𝗱! Also, check out my health AI blog here: https://lnkd.in/g3nrQFxW
-
Probably, one of the largest collaborative efforts in biotech, since the Human Genome Project: the Human Cell Atlas has arrived! 🧬 I think the Human Cell Atlas (HCA) is a pretty monumental leap in systems biology, an international effort involving 3,600 researchers from 102 countries, has released its first draft atlas of human cells. This isn’t just another dataset—this is the blueprint of human biology, built cell by cell, tissue by tissue, organ by organ. The HCA integrated data from 62 million cells, sourced from 9,100 donors, spanning every stage of human development—embryonic to adult. Researchers organized their work into 18 Biological Networks, focusing on key organs like the lung, nervous system, and eye. Some of the tools like single-cell RNA sequencing, spatial transcriptomics, and multi-omics were combined to profile and map cells with unprecedented precision. Notably, Google provided essential cloud infrastructure and AI tools like scTab (for annotation) and SCimilarity (for cell similarity searches), helping researchers handle vast and complex datasets efficiently. It is also important that local scientists and the HCA Ethics Working Group put efforts to make sure data represented populations globally, prioritizing equity and open access. Now, how can we use it, practically speaking? Here I picked some of the key aspects that might be very useful for the biotech community: ✅ Precise Target Discovery: Pinpoint disease-specific cell types and biomarkers to create highly targeted therapies. ✅ Better Disease Models: Build realistic organoids and in vitro models informed by detailed cell maps for accurate drug testing. ✅ Personalized Medicine: Utilize data from diverse populations to design therapies tailored to genetic and environmental variations. ✅ Safer Drugs: Analyze tissue-specific metabolism to predict and avoid adverse drug effects. ✅ AI-Driven Insights: Tap into machine-learning tools like PopV and SCimilarity to accelerate discovery and refine findings. I believe, the Atlas could be a playing ground for other AI tools and new workflows! ✅ Early Diagnosis: Identify subtle gene expression changes for early detection of diseases like cancer or neurodegenerative disorders. If you're in biotech, drug discovery, or systems biology, this resource is now open and available—check it out! Link in the comments 👇 Image source: Springer Nature
-
🚀 GROMACS Workflow — From Structure to Scientific Insight🖥️💻 Molecular dynamics isn’t just running simulations — it’s about engineering reproducible, physically meaningful systems. Here’s a distilled breakdown of the full workflow with practical insights 👇 🔬 1. System Preparation — Where accuracy begins Start with a clean PDB structure Choose an appropriate force field (AMBER/CHARMM/OPLS) — this decision directly impacts your results Define simulation box (⚠️ dodecahedron saves computational cost vs cubic) Solvation + ion addition ensures physiological realism 💡 Real insight: Many beginners overlook ion concentration — only add it when mimicking real biological conditions, not blindly. ⚙️ 2. Equilibration → Production MD — Stabilize before you simulate Energy Minimization: removes steric clashes NVT (constant T): stabilizes temperature NPT (constant P): stabilizes density & pressure Production Run: actual data generation 💡 Real insight: If your system isn’t stable in NVT/NPT, your production data is scientifically unreliable — no shortcuts here. 📊 3. Trajectory Analysis — Data → Meaning Key metrics: RMSD → structural stability RMSF → residue flexibility Radius of gyration → compactness H-bonds → interaction strength SASA → solvent exposure 💡 Real insight: Don’t just plot graphs — interpret trends biologically (e.g., RMSD plateau = stable conformation). 🧠 4. Advanced Analysis — Where research gets publishable PCA → dominant motions in the system Free Energy (MM/PBSA) → binding affinity Clustering → dominant conformations 💡 Real insight: PCA often reveals hidden conformational states that RMSD alone cannot detect. 🎥 5. Visualization & Reporting — Communicate like a scientist Tools: VMD, PyMOL, Chimera Generate trajectories, publication figures, movies 💡 Real insight: A well-visualized result often communicates better than raw data — this is what reviewers notice. 🔥 Key Takeaway: GROMACS is not just a tool — it’s a pipeline of decisions. Each step (force field, box type, equilibration strategy) shapes your final scientific conclusion. 💬 If you're stepping into computational biology / bioinformatics / MD simulations, mastering this workflow gives you a serious edge. Here are the attached Link 🔗 https://lnkd.in/g45E26RC 👉 Follow me for more insights like this — breaking down complex research workflows into clarity. #GROMACS #MolecularDynamics #Bioinformatics #ComputationalBiology #ResearchSkills #LifeSciences #ScientificComputing #DataAnalysis #Biotechnology #STEM #GraduateStudies #PhDJourney
-
In a recent The Washington Post conversation, I described why I left Coursera in 2016 to return to the convergence of AI and biology: the realization that we were at the nexus of two tidal waves: biological data at scale, enabled via technologies like single-cell RNA sequencing, human-relevant cell differentiation, and super-resolution microscopy; and modern AI’s ability to see patterns a human will never discern. The United States Senate Bipartisan Commission on Biotechnology has now articulated a critical point: that biological data should be treated as a strategic national resource. Their final report proposes a Web of Biological Data (WOBD) — a unified, AI-ready infrastructure for accessing biological datasets — paired with NIST standards for AI-readiness and a network of automated cloud labs. This is exactly right, and the timing matters. A national, AI-ready WOBD could become the shared substrate on which the entire field builds. Large language models emerged not from clever architecture alone but from training on an internet's worth of text and images. Biological AI models need their equivalent. The bottleneck to better medicines is not chemistry or molecular biology — it is high-quality, interoperable biological data at the scale modern ML requires. AI's potential in biology will be truly unlocked at population scale on the human side combined with massive amounts of human-relevant cellular data — and that data doesn't yet exist publicly at the quality AI requires. The Commission's cloud labs proposal could deliver a step function in data generation. Automated, high-throughput data generation is genuinely transformative — but also one of the hardest engineering challenges in the field. Building automation that produces data at scale is hard; ensuring the AI learns signal rather than artifact is harder still. At insitro we've been at this for eight years and have made real progress. That experience has taught us this capability takes time to develop and compound, which is precisely why national investment needs to start now. If there is one priority I would emphasize above all others, it is human data at scale. In vitro systems are invaluable — they enable causal interventions that population genetics alone cannot provide. But human data is the ultimate ground truth. Genetics is causal by definition: variation that arose before disease did. The UK Biobank's 500,000 deeply phenotyped individuals showed what becomes possible at that threshold — data so rich that machine learning sees things no human ever could. The WOBD's highest-leverage investment would be dramatically expanding the scale and accessibility of human biological data for U.S. researchers — something the U.K. and other nations have done better than we have. We are at a once-in-a-generation inflection point. The science is ready. The AI is ready. The Senate Commission just correctly named the remaining bottleneck. I hope Congress acts.
-
Analyzing MD simulation data can be challenging when you're starting in protein modeling. In my most recent workshop, I showed clients how to streamline this process using a Google Colab notebook that I called MD_quick_plot. A tool designed to simplify MD simulation data analysis so you can spend more time formulating hypotheses rather than dealing with complex workflows to generate plots. All you need is a structural file (PDB, Gro) and a trajectory file (xtc, trr, or dcd), and that is it!! MD_quick_plot allows you to readily visualize: -RMSD (Root Mean Square Deviation) -RMSF (Root Mean Square Fluctuation) -Radius of Gyration (Rg) -Free Energy Landscape (FEL) -MM Binding Energy (Electrostatic + VdW) -Protein-Ligand Distance Analysis If you want to check it out. In the comments, I'll leave the link to the GitHub repository where you can access the Google Colab notebook. Feedback is always welcome!!! Follow me, Omar Arias-Gaguancela, PhD, and SciLearningWorkshops for more content like this!!!
-
Where #AlphaFold prediction meets cryo-EM density A recently deposited cryo-EM structure of a lipid export complex contains a small tracing issue and a 4-residue register shift in a transmembrane helix crucial for complex formation. Four residues are almost exactly one turn of an α-helix (~3.6 residues per turn). The mostly hydrophobic interface still looks perfectly reasonable. And yet, it’s wrong. After correcting the register and tracing, map fit improves, a plausible Phe–Met contact appears, and a proline moves from a helix cap into the helix core gaining functional significance. Interestingly, in this case the AlphaFold predicts the correct interface. However, details vary between repeated predictions while showing similar confidence scores. This illustrates that pLDDT, ipTM, and PAE are not designed to validate residue-level details in small interfaces. AI tools provide strong structural priors. Cryo-EM provides often weak but real experimental constraints. Tools that aggregate subtle map signals (like #checkMySequence) or extract residue-level contacts from weak evolutionary couplings (like #gapTrick) don’t replace human interpretation. They amplify weak evidence. The future of structural biology is not AI versus experiment. It’s iterative cross-validation between them.
-
In light of recent work on LLM-based single-cell annotation, I created a R function for you that allows you to integrate this into your workflow, and makes explicit how it works, so you can be empowered to develop things like this on your end without relying on high-level interfaces... The LLM: I use OpenRouter, which gives me API access to the likes of GPT's, Claude, and DeepSeek without being locked into one vendor. You can use any of these if you use my tool. My function converts each cluster's output from Seurat's FindAllMarkers() into a string, which gets combined with a prompt fed into the LLM, per cluster. The output is a vector of annotated cell populations. Results: In this experiment, I used Claude 3.5 Sonnet on the back end. This tool was able to annotate the PBMC 3k dataset accurately, with errors involving depth of classification (eg. stopping at CD4 T, and not choosing naive or memory). Running the tool multiple times revealed wording changes (eg. CD4 T vs CD4+ T) but not changes in population guess. Discussion: Complexity of data: The PBMC datasets are simple and well-trodden. It is likely that LLM use will trip up in weird ways when we start looking at more complex data, like developmental trajectories or cancer. Sophistication of model: Claude 3.5 Sonnet is a relatively good model at the time of writing, but we note that if this function trips up on more complex data, the user can switch to DeepSeek R1 or any of the other reasoning models for testing. Accuracy will likely get better as the models become more sophisticated. A future direction here is fine-tuning a model or using a model fine-tuned for the task of cell annotation (see my posts on foundation models). Prompt engineering: The prompt is relatively straightforward, and there is room to play around with the prompt itself. One simple example might be to provide a document of examples of annotated cell types and what genes they express, directly as a pre-prompt. Such a document is increasingly more possible now, given all the single-cell "atlases" that are being constructed. Try it yourself: Use my tool (or similar ones). To use it, just get an OpenRouter API key. The rest is simply copying and pasting a block of code. Battle test it on your "real world data." Let me know where the model trips up. This will allow me and others working on similar things to figure out how to improve these things down the line. Doing similar project on your end? Reach out. Plenty of people are talking about LLMs but few are actually doing work on them, and I would like to know who you are. The R markdown with the respective code is linked in the comments. Have fun with it.
Explore categories
- Hospitality & Tourism
- Productivity
- Finance
- Soft Skills & Emotional Intelligence
- Project Management
- Education
- Technology
- Leadership
- Ecommerce
- User Experience
- Recruitment & HR
- Customer Experience
- Real Estate
- Marketing
- Sales
- Retail & Merchandising
- Supply Chain Management
- Future Of Work
- Consulting
- Writing
- Economics
- Artificial Intelligence
- Employee Experience
- Healthcare
- Workplace Trends
- Fundraising
- Networking
- Corporate Social Responsibility
- Negotiation
- Communication
- Engineering
- Career
- Business Strategy
- Change Management
- Organizational Culture
- Design
- Innovation
- Event Planning
- Training & Development