How Undergraduate Research and Open-Source Software Are Transforming Genetic Visualization
A new open-source software tool called pygenoscape is giving biologists a powerful way to visualize how genetic variation changes across geographic space, answering fundamental questions about population isolation and genetic diversity that researchers have struggled with for decades. The software, developed by Assistant Professor of Biology Andrew Davinack and undergraduate Rylie Seaberg at Wheaton College Massachusetts, transforms thousands of DNA sequences into interactive three-dimensional landscapes that reveal genetic patterns invisible in traditional data analysis.
What Problem Does Pygenoscape Solve?
For years, biologists have faced a frustrating limitation: existing software couldn't handle large-scale genetic datasets or provide intuitive visualizations of how populations differ genetically across landscapes. The most comparable tool, Alleles in Space, was released in 2005 and only works on Windows computers, contains bugs, and struggles with large datasets. Seaberg encountered this problem firsthand while working on her honors thesis, assembling and analyzing thousands of genetic sequences from human and veterinary parasites.
Pygenoscape solves this by taking genetic distance data and overlaying it in three-dimensional space, creating visual landscapes where peaks and valleys represent genetic variation. A completely flat landscape indicates genetically similar populations, while sharp peaks show regions of greater genetic diversity. This intuitive visualization helps researchers answer critical questions: Are populations becoming genetically isolated simply because they live far apart, or are physical barriers like rivers, mountains, or dams preventing animals from interbreeding ?
"If you want to study how genetically related gray wolf populations are in Yellowstone National Park, for example, you can extract DNA from 200 animals and compare their genetic distances. This software takes those data points and overlays them in a three-dimensional space. A completely flat landscape means the wolves are genetically similar, but an uneven landscape interrupted by sharp peaks indicate regions where there is much greater genetic variation," explained Andrew Davinack, Assistant Professor of Biology at Wheaton College Massachusetts.
Andrew Davinack, Assistant Professor of Biology at Wheaton College Massachusetts
Why Does This Matter for Modern Biology Research?
Unlike earlier programs, pygenoscape is designed to work with both traditional DNA sequence data and modern genome-scale datasets containing thousands of genetic markers. The software is also freely available, making it accessible to researchers worldwide. It arrives at a time when biological researchers are increasingly turning to Python, a general-purpose programming language, for large-scale data analysis instead of relying solely on R, which has historically dominated the biological sciences.
The tool has practical applications beyond academic research. In conservation biology, pygenoscape can help determine whether removing dams or other infrastructure will restore genetic connectivity among aquatic species. These visualizations give conservation managers concrete data to decide whether particular conservation strategies should move forward.
How to Use Pygenoscape for Genetic Research
- Data Input: Load thousands of genetic sequences from databases or your own research into the software with a single line of Python code.
- Visualization Generation: The software automatically creates three-dimensional landscapes showing how genetic variation changes across geographic space without requiring manual data manipulation.
- Pattern Recognition: Identify peaks and valleys in the landscape to understand genetic isolation, population connectivity, and the effects of geographic or physical barriers on species.
- Conservation Planning: Use the visualizations to inform decisions about infrastructure removal, habitat restoration, or species management strategies.
The software was tested on datasets ranging from invasive bumblebees to marine invertebrates before Davinack and Seaberg published their findings in Bioinformatics Advances. Seaberg, who compiled and processed thousands of sequences for her honors thesis, became the first major tester of pygenoscape, helping debug and validate the software.
What Does This Tell Us About the Future of Bioinformatics?
The project demonstrates how undergraduate research can contribute meaningfully to scientific advances and highlights the growing demand for bioinformaticians. As biology becomes increasingly data-intensive, professionals who combine computational skills with scientific knowledge are in high demand. Bioinformaticians can work at nonprofits, hospitals, sequencing facilities, biotech startups, and organizations with specific missions like the National Cancer Institute.
"There's a need for more people who know how to work with a lot of data, especially with the rise of AI," noted Andrew Davinack.
Andrew Davinack, Assistant Professor of Biology at Wheaton College Massachusetts
Wheaton College's interdisciplinary bioinformatics program combines biology, mathematics, computer science, biological chemistry, and environmental science, allowing students to design their own research paths rather than following a single prescribed curriculum. This approach encourages intellectual curiosity and provides students with unique experiences, such as helping design bioinformatics software packages through honors-level research.
For Seaberg, the experience confirmed her career direction. After completing the project, she realized she enjoyed analyzing data patterns more than studying specific genes and genomics. She will begin pursuing a master's degree in data science at Boston University this fall, building on the mentorship and research experience she gained at Wheaton.