Loading…

Registration opens August 15th at www.idwsds.org.
Type: Plenary Talk clear filter
arrow_back View All Dates
Tuesday, October 6
 

00:30 UTC

S102 - Gini-Weighted Risk Control in Adaptive Generative Augmentation
Tuesday October 6, 2026 00:30 - 01:00 UTC
Generative data augmentation is widely used to address class imbalance by enriching minority classes with synthetic samples. Existing approaches typically employ a fixed augmentation strength across all classes, ignoring differences in class imbalance, data structure, and generative quality. We propose an adaptive augmentation framework that determines class-specific augmentation strengths using both class proportions and structural information measured through Gini correlation. This strategy allocates more synthetic data to underrepresented and weakly structured classes while limiting augmentation for well-represented classes.

We develop a theoretical framework showing that the resulting excess risk is controlled by a weighted combination of class-conditional Wasserstein discrepancies and Gini-based structural factors. We further establish consistency results and demonstrate that adaptive augmentation provides tighter control of risk distortion than fixed augmentation schemes.

Experiments on imbalanced classification datasets show consistent improvements in minority-class recall and macro-F1 performance, while empirical results closely match theoretical predictions. The proposed framework provides a principled, structure-aware foundation for generative data augmentation.
Speakers
avatar for Chathurika Abeykoon

Chathurika Abeykoon

Assistant professor of Mathematics and Statistics, Rhodes College
Chathurika Abeykoon is an Assistant Professor of Statistics at Rhodes College, Memphis, TN. Dr. Abeykoon received her Ph.D. in Mathematics with a concentration in Statistics from the University of Mississippi in 2023. Her research lies at the intersection of Mathematics, Statistics... Read More →
Organizers
avatar for Chathurika Abeykoon

Chathurika Abeykoon

Assistant professor of Mathematics and Statistics, Rhodes College
Chathurika Abeykoon is an Assistant Professor of Statistics at Rhodes College, Memphis, TN. Dr. Abeykoon received her Ph.D. in Mathematics with a concentration in Statistics from the University of Mississippi in 2023. Her research lies at the intersection of Mathematics, Statistics... Read More →
Tuesday October 6, 2026 00:30 - 01:00 UTC
Zoom Room #1

00:30 UTC

S201- Vision-Language Models and the Harms of Covert Sexualization
Tuesday October 6, 2026 00:30 - 01:00 UTC
Vision-language models (VLMs) increasingly mediate how bodies are described, moderated, and rendered in online spaces through content moderation, AI-generated images and videos, and descriptions of real bodies. Prior work establishes that VLMs sexually objectify bodies that are partially clothed more than fully clothed ones, and that plus-size bodies — particularly those belonging to women and AFAB persons — are disproportionately censored on social media platforms for being “inappropriate.” However, scholars lack the tools to detect when a model *objectifies* a body, rather than merely describing it, and whether this behavior differs across body *shapes*, rather than just body sizes. This talk describes a study in which we investigate these phenomena using swimwear try-on images matched across body morphology, comparing high-contrast silhouettes (i.e., a smaller waist relative to hips and bust) against lower-contrast silhouettes under identical prompting conditions. We ask whether VLMs describe these bodies differently despite equivalent context and whether unwanted objectification or descriptive drift disproportionately burdens those with high-contrast bodies, with a focus on covert (i.e., safety guardrail compliant) sexualization. By making the harms of these phenomena visible and measurable, this work gives researchers and auditors the ability to hold systems accountable when they objectify or censor women and AFAB persons based on their appearance.
Speakers
avatar for Maimuna Majumder

Maimuna Majumder

Harvard Medical School & Boston Children's Hospital
Dr. Maimuna (Maia) Majumder (she/they), MPH, PhD (MIT '18) is an Assistant Professor and Inaugural Peter Szolovits Distinguished Scholar in the Computational Health Informatics Program at Harvard Medical School and Boston Children’s Hospital. She is a computational epidemiologist... Read More →
Organizers
avatar for Maimuna Majumder

Maimuna Majumder

Harvard Medical School & Boston Children's Hospital
Dr. Maimuna (Maia) Majumder (she/they), MPH, PhD (MIT '18) is an Assistant Professor and Inaugural Peter Szolovits Distinguished Scholar in the Computational Health Informatics Program at Harvard Medical School and Boston Children’s Hospital. She is a computational epidemiologist... Read More →
Tuesday October 6, 2026 00:30 - 01:00 UTC
Zoom Room #2

00:30 UTC

S301 - Why and When Targeted Validation Sampling Offers Statistical Efficiency Gains: A Case Study on Healthy Food Access and Disease Outcomes
Tuesday October 6, 2026 00:30 - 01:00 UTC
Quantifying neighborhood food environments and understanding their relationships with residents’ health is a public health priority. Using simple, error-prone food access metrics (like the shortest straight-line routes to healthy food stores) introduces measurement error and biases downstream statistical models, but measuring the more-accurate, map-based ones (like shortest driving routes) for entire studies is often implausible. Fortunately, adopting a two-phase design can harness the best of both metrics by combining the error-prone access measures for the entire study and the more-accurate ones for a chosen subset in a partial validation study. This validated subset can be strategically chosen to not only reduce bias but further improve efficiency when modeling relationships between health and the food environment. Technically, any information that is fully available for all neighborhoods can guide the validation sampling strategy. One such promising design paired stratification with Neyman allocation and sampled based on the influence function within each stratum. Using simulations and data for the Piedmont Triad Region of North Carolina, various validation sampling designs were evaluated to quantify the associations of diabetes count and obesity prevalence with neighborhood-level access to healthy foods, fitting two separate Poisson regression models, one for each outcome. We assess which design suits each model, and whether any is robust across settings.
Speakers
avatar for L. Ishara Wijayaratne

L. Ishara Wijayaratne

Department of Statistics, Wake Forest University
I am a graduate student in Statistics at Wake Forest University. Originally from Sri Lanka, I am keen on giving back to the community that has helped change my life so considerably. I am passionate about biostatistics, and my abstract submission addresses a pressing public health... Read More →
Organizers
avatar for L. Ishara Wijayaratne

L. Ishara Wijayaratne

Department of Statistics, Wake Forest University
I am a graduate student in Statistics at Wake Forest University. Originally from Sri Lanka, I am keen on giving back to the community that has helped change my life so considerably. I am passionate about biostatistics, and my abstract submission addresses a pressing public health... Read More →
Tuesday October 6, 2026 00:30 - 01:00 UTC
Zoom Room #3

07:00 UTC

S107 - Absorbing Markov Chain Parameter Estimation Under Data Scarcity: A Comparative Study of Analytical and Monte Carlo Methods in Neonatal Care
Tuesday October 6, 2026 07:00 - 07:30 UTC
Neonatal mortality remains a major global public health challenge, with an estimated 6,200 newborns dying daily, mostly in settings where patient records are scarce. For data-scarce neonatal units, a key question arises: when transition data is limited, does it matter whether Analytical Estimation or Monte Carlo Simulation is used to model patient outcomes?

This study addresses that question using a neonatal dataset of 6,000 daily state transitions across 2,000 patients in Ghana. An absorbing Markov chain with four states: Hospital Admission, Neonatal Intensive Care Unit (NICU), Recovered, and Death was used. Both methods were evaluated across six data levels with 500 independent replications using bias, variance, standard deviation, and mean squared error.

Both methods perform similarly with adequate data and degrade equivalently under data scarcity because they share the same estimated transition matrix. Expected time to absorption is more sensitive to limited data than absorption probabilities, and 500 observed transitions emerge as the minimum reliable threshold. These findings provide an evidence-based data standard for resource-constrained healthcare systems.

Keywords: Absorbing Markov Chain, Monte Carlo Simulation, Analytical Estimation, Absorption Probability, Expected Time to Absorption, Sensitivity Analysis.

Authors: Emmanuella Frimpong, Dr. Irene Kafui Vorsah Amponsah, PhD (Visiting Lecturer at Ohio University)
Speakers
avatar for Emmanuella Frimpong

Emmanuella Frimpong

Miami University, Oxford, Ohio
Ms. Emmanuella Frimpong is a graduate of the African Institute for Mathematical Sciences (AIMS) Ghana, where she completed her Master of Science in Mathematics, undertaking research on the topic “Comparing Monte Carlo Simulation and Analytical Estimation Methods for Absorbing Markov... Read More →
Organizers
avatar for Emmanuella Frimpong

Emmanuella Frimpong

Miami University, Oxford, Ohio
Ms. Emmanuella Frimpong is a graduate of the African Institute for Mathematical Sciences (AIMS) Ghana, where she completed her Master of Science in Mathematics, undertaking research on the topic “Comparing Monte Carlo Simulation and Analytical Estimation Methods for Absorbing Markov... Read More →
Tuesday October 6, 2026 07:00 - 07:30 UTC
Zoom Room #1

07:00 UTC

S204 - Creating Global Connections Through Early Statistical Leadership in Rare Disease Trials
Tuesday October 6, 2026 07:00 - 07:30 UTC
In complex clinical research, collaboration begins well before database lock or final analysis. It begins when statisticians are involved early enough in protocol development to shape what is measurable, sustainable, and meaningful for patients. This presentation uses a rare pediatric dermatology trial as a case study to show how early statistical input can prevent avoidable missing data before the first participant is enrolled.

The original design was scientifically ambitious but operationally burdensome, with frequent in-clinic visits, narrow assessment windows, and substantial caregiver burden. Statistically, this created a foreseeable risk of informative missingness, loss to follow-up, and reduced interpretability in a small and vulnerable population. Early statistical leadership helped reframe the design around feasibility as well as rigor.

The talk will illustrate how collaboration among statisticians, clinicians, programmers, operations teams, and digital partners can reduce this risk. Examples include identifying burden-sensitive endpoints, anticipating missing-data pathways, and considering remote ePRO, image capture, or AI-supported skin assessment to replace selected in-person visits without compromising data quality.

The broader message is that statistics is not only an analysis discipline; it is also a design and connecting discipline.
Speakers
avatar for Shrutimita Pokhariyal

Shrutimita Pokhariyal

PHASTAR
Shrutimita Pokhariyal is an experienced biostatistics professional with over 14 years in clinical research, with expertise spanning multiple phases of clinical trials. Her work includes statistical strategy, study design, protocol and SAP development, and collaboration across cross-functional... Read More →
Chairs/Hosts
avatar for Shrutimita Pokhariyal

Shrutimita Pokhariyal

PHASTAR
Shrutimita Pokhariyal is an experienced biostatistics professional with over 14 years in clinical research, with expertise spanning multiple phases of clinical trials. Her work includes statistical strategy, study design, protocol and SAP development, and collaboration across cross-functional... Read More →
Organizers
avatar for Shrutimita Pokhariyal

Shrutimita Pokhariyal

PHASTAR
Shrutimita Pokhariyal is an experienced biostatistics professional with over 14 years in clinical research, with expertise spanning multiple phases of clinical trials. Her work includes statistical strategy, study design, protocol and SAP development, and collaboration across cross-functional... Read More →
Tuesday October 6, 2026 07:00 - 07:30 UTC
Zoom Room #2

07:30 UTC

S108 - Explainable AI in Health Technology
Tuesday October 6, 2026 07:30 - 08:00 UTC
The rapid deployment of machine learning in healthcare has exposed a fundamental tension between model performance and clinical utility. This talk addresses the urgent need for Explainability (XAI), moving beyond the "black box" paradigm to ensure that AI systems are not only accurate but also transparent, accountable, and trustworthy. In the high-stakes environment of medical diagnostics, understanding an algorithm’s pathway is essential for verifying clinical reliability and meeting the rigorous standards demanded by both clinicians and regulators.
A central theme of this talk is the transition from surface-level performance metrics to deep reproducibility and validation. We will examine how traditional accuracy scores can sometimes be deceptive and how to combat this, compare XAI with true interpretability, and tackle the challenge of the inheritance of bias. This session outlines proactive mitigation strategies, including rigorous data quality checking, subgroup analysis, and continuous fairness assessments throughout the model lifecycle.
Finally, the talk will navigate the evolving regulatory landscape, specifically the EU AI Act.
Speakers
avatar for Autumn Johnson

Autumn Johnson

University of Galway
My name is Autumn Johnson, and I’m a postdoctoral researcher in Statistics at the University of Galway. My postdoctoral work has focused on health technology advancements using complex statistical methods, machine learning, and AI. My PhD is from University College Cork, where I... Read More →
Organizers
avatar for Autumn Johnson

Autumn Johnson

University of Galway
My name is Autumn Johnson, and I’m a postdoctoral researcher in Statistics at the University of Galway. My postdoctoral work has focused on health technology advancements using complex statistical methods, machine learning, and AI. My PhD is from University College Cork, where I... Read More →
Tuesday October 6, 2026 07:30 - 08:00 UTC
Zoom Room #1

09:00 UTC

S206 - Beyond Simulation: Benchmarking for Method Comparison in Statistics and Machine Learning
Tuesday October 6, 2026 09:00 - 09:30 UTC
When initiating a statistical or machine learning analysis, one of the first and most consequential questions is: which method should be used for a given dataset? Traditional approaches for comparing methods include theoretical derivations and data simulation studies, which provide insight into method performance under controlled conditions. However, these approaches may not fully reflect the complexity of real-world data. Benchmarking—systematic comparison of methods across many real datasets—offers a complementary approach that can improve generalizability and provide practical guidance. Despite its common use in computer science, benchmarking remains underutilized in statistical methodology as new methods continue to emerge.

In this talk, I will discuss benchmarking in the context of statistical and machine learning research and contrast it with theory and simulation. I will outline key principles for conducting rigorous benchmarking studies and illustrate them using two case studies: benchmarking random forest variable selection methods for categorical and continuous outcomes and comparing methods for time-to-event data using the mlr3 framework. Together, these examples demonstrate how benchmarking can enhance scientific rigor, provide practical guidance for method selection, and support more transparent and reproducible methodological research.
Speakers
avatar for Jaime Speiser

Jaime Speiser

Associate Professor of Biostatistics and Data Science, Wake Forest University School of Medicine
Dr. Speiser is a biostatistician focused on prediction modeling with applications in medicine. Her work involves developing novel machine learning methodology for prediction modeling, providing guidance on best practices for developing prediction models, and collaborating with medical... Read More →
Organizers
avatar for Jaime Speiser

Jaime Speiser

Associate Professor of Biostatistics and Data Science, Wake Forest University School of Medicine
Dr. Speiser is a biostatistician focused on prediction modeling with applications in medicine. Her work involves developing novel machine learning methodology for prediction modeling, providing guidance on best practices for developing prediction models, and collaborating with medical... Read More →
Tuesday October 6, 2026 09:00 - 09:30 UTC
Zoom Room #2

09:30 UTC

S207 - On similarity-based models for Bayesian disease mapping
Tuesday October 6, 2026 09:30 - 10:00 UTC
In the 1970s, the conditionally formulated Gaussian Markov random field (GMRF), known as the conditional autoregressive (CAR) model, was introduced in line with Tobler’s first law of geography: "everything is related to everything else, but near things are more related than distant things." The CAR model uses W, the well-known adjacency matrix that encodes the neighbourhood structure of a spatial lattice (e.g., two areas are neighbours if they share a common border).
Since its introduction, the CAR model has undergone many adaptations, with numerous adaptive models proposed in the literature. However, almost all of these models (to the best of our knowledge, all except ours) still adhere to Tobler’s law. Yet, much of the data collected today – often aggregated at the areal level for confidentiality or other reasons – does not necessarily follow this law.
In this talk, we will show that any extra information representing causes of or correlated with the phenomenon of interest can be used to define a similarity structure, rather than relying solely on geographical neighbourhoods. Using simulated data, we illustrate that similarity-based structures can be more effective than traditional neighbourhood-based structures for smoothing both local and global risks. We will show that the correct identification of high- and low-risk areas, crucial for public health planning and resource allocation, is better achieved when the similarity-based structure is used.
Speakers
avatar for Helena Baptista

Helena Baptista

Management Information Centre (MagIC), NOVA Information Management School (NOVA IMS), Universidade Nova de Lisboa, Campus de Campolide, 1070-312, Lisboa, Portugal
Helena Baptista is a highly experienced statistician, researcher, and educator, specializing in applied statistics, forecasting, and time series analysis. She has over 25 years of experience in the pharmaceutical industry, finance, and academia, with a strong background in statistical... Read More →
Organizers
avatar for Helena Baptista

Helena Baptista

Management Information Centre (MagIC), NOVA Information Management School (NOVA IMS), Universidade Nova de Lisboa, Campus de Campolide, 1070-312, Lisboa, Portugal
Helena Baptista is a highly experienced statistician, researcher, and educator, specializing in applied statistics, forecasting, and time series analysis. She has over 25 years of experience in the pharmaceutical industry, finance, and academia, with a strong background in statistical... Read More →
Tuesday October 6, 2026 09:30 - 10:00 UTC
Zoom Room #2

10:00 UTC

S208 - Optimal Model selection for incidence of Birth Asphyxia: NICU Centers in Greater Accra Region.
Tuesday October 6, 2026 10:00 - 10:30 UTC
Abstract
Birth asphyxia remains a major contributor to neonatal morbidity and mortality in low- and middle-income countries, particularly in sub-Saharan Africa. This study investigated the determinants of birth asphyxia among newborns using a Quasi-Poisson regression model to account for overdispersion in the count data. Secondary data comprising neonatal and maternal records were analysed using descriptive statistics, correlation analysis, and inferential modelling. The Quasi-Poisson model selected after diagnostic assessment confirmed overdispersion in the response variable, making it more appropriate than the standard Poisson model. The Quasi-Poisson regression model for predicting birth asphyxia is given as : Birth Asphyxia = 1.742 +0.385(Birth Weight) + 0.012(Mothers Age)− 0.088(Gestational Age) + 0.0401(Sex)+ 0.301(Mode of Delivery) + 0.158(Parity)+ 0.067(SURVIVE)
The results showed that birth weight, gestational age, mode of delivery, maternal age, parity, and fetal presentation were significant predictors of birth asphyxia. Specifically, lower birth weight and shorter gestational age were associated with a higher incidence of birth asphyxia, while caesarean delivery and abnormal fetal presentation increased the likelihood of adverse birth outcomes.
Keywords: Birth asphyxia, Quasi-Poisson regression, Gestational age, Neonatal outcomes, Maternal health, Ghana.
Authors:
1. Selina Dadzie (Mphil)
2. Irene Kafui Vorsah Amponsah (PhD)
Speakers
avatar for Selina Dadzie

Selina Dadzie

Student
Bio
Miss Selina Dadzie is a graduate student in Statistics at the University of Cape Coast, Ghana, where she is pursuing a Master of Philosophy (MPhil) in Statistics. Her research focuses on “Optimal Model Selection for Incidence of Birth Asphyxia: NICU Centers in Accra”, with particular... Read More →
Organizers
avatar for Selina Dadzie

Selina Dadzie

Student
Bio
Miss Selina Dadzie is a graduate student in Statistics at the University of Cape Coast, Ghana, where she is pursuing a Master of Philosophy (MPhil) in Statistics. Her research focuses on “Optimal Model Selection for Incidence of Birth Asphyxia: NICU Centers in Accra”, with particular... Read More →
Tuesday October 6, 2026 10:00 - 10:30 UTC
Zoom Room #2

10:30 UTC

S209 - Temporal and Regional Variations of Effects of Daily Temperature on Annual Precipitation: A Functional Mixed Effect Model Approach
Tuesday October 6, 2026 10:30 - 11:00 UTC
"Given the increasing threat of global warming, it is important to understand not only
global weather patterns but also regional variations throughout countries. Functional
Linear Mixed-effects Model (FLMM), an emerging statistical tool, provides a comprehensive
framework for analyzing functional data (e.g., data viewed as a function
or curve) with repeated observations, allowing researchers to identify patterns and relationships
in the data. This study applies the functional linear mixed-effects model
(FLMM) to recognize the effects of temporal and regional variations of short-time
weather projection (monthly precipitation on daily temperature). We deployed FLMM
using the daily temperature and monthly precipitation of nine weather stations(Dhaka,
Chattogram (Patenga), Chattogram (Ambagan), Rajshahi, Khulna, Barisal, Sylhet,
Rangpur and Mymensingh) of Bangladesh where each station shares the common
population effects with their individual scalar covariate effects along with same slope
functions. To estimate variance parameters and fixed-effects and random effects, the
REML-based EM algorithm proposed by has been applied. Empirical
results show significant differences in the effects of monthly precipitation on daily
temperature among regions of Bangladesh. We anticipated that FLMM is an emerging
model that can reveal the effects of monthly precipitation and the pace of daily
temperature fluctuations over time."
Speakers
avatar for Munniara Yesmin Munni

Munniara Yesmin Munni

PME Officer
Munniara Yesmin Munni is a statistician and Monitoring, Evaluation, and Learning (MEL) professional with a Master’s degree in Statistics from Jahangirnagar University, Bangladesh. She currently works with the Christian Commission for Development in Bangladesh (CCDB). Her research... Read More →
Organizers
avatar for Munniara Yesmin Munni

Munniara Yesmin Munni

PME Officer
Munniara Yesmin Munni is a statistician and Monitoring, Evaluation, and Learning (MEL) professional with a Master’s degree in Statistics from Jahangirnagar University, Bangladesh. She currently works with the Christian Commission for Development in Bangladesh (CCDB). Her research... Read More →
Tuesday October 6, 2026 10:30 - 11:00 UTC
Zoom Room #2

12:00 UTC

S112 - Predicting the Right Treatment for the Right Patient: An AI-Powered Decision Support Framework Based on Predicted Individual Treatment Effects
Tuesday October 6, 2026 12:00 - 12:30 UTC
Medical decisions—ranging from diagnosis to treatment selection—are inherently uncertain. Clinicians often rely on heuristic, experience-driven processes to integrate heterogeneous data. In this context, Predicted Individual Treatment Effects (PITE) offer a principled statistical framework to quantify how much a specific patient benefits from one treatment over another.

Advances in computational systems now enable machines to identify complex patterns within large datasets, facilitating a shift toward data-driven, individualized healthcare. This work explores using PITE to support clinical decision-making across diverse diseases and contexts. We address critical questions: which AI methods suit specific clinical datasets, how PITE should be validated, and how outcome complexity affects tool reliability.

We demonstrate PITE-based models in various disease settings, each posing unique methodological challenges. Our results show that even under real-world conditions—such as missing data—predictive models maintain interpretability and generate estimates that support clinicians. Notably, our findings highlight that internal validation is insufficient; external validation is essential for robust predictions.

Ultimately, effective PITE-based support requires more than modeling. It demands an adaptive, continuously learning system integrating data management, modeling strategies, regulatory-grade interpretability, and ongoing validation to translate evidence into precise.
Speakers
avatar for Pamela Solano

Pamela Solano

PhD Researcher, Faculty of Computer Science and Data Science, Regensburg University
I am Pamela Solano, a statistician and researcher at the University of Regensburg, Germany. Since 2014, I have worked as a biostatistician. Following my PhD in 2018, my research focus toward statistical modeling approaches with direct societal relevance. I began working in environmental... Read More →
Organizers
avatar for Pamela Solano

Pamela Solano

PhD Researcher, Faculty of Computer Science and Data Science, Regensburg University
I am Pamela Solano, a statistician and researcher at the University of Regensburg, Germany. Since 2014, I have worked as a biostatistician. Following my PhD in 2018, my research focus toward statistical modeling approaches with direct societal relevance. I began working in environmental... Read More →
Tuesday October 6, 2026 12:00 - 12:30 UTC
Zoom Room #1

12:30 UTC

S113 - Navigating Data Sharing in Medical Research
Tuesday October 6, 2026 12:30 - 13:00 UTC
Open science is pivotal in advancing medical research by promoting accessibility and collaboration among researchers globally, and biostatisticians play a critical role in supporting these efforts. Shared datasets and software code, particularly those developed using time-intensive algorithms, are fundamental for fostering replicability and improving research efficiency for other scientists. Open access publications accompanied by publicly available data and code also help reduce disparities in knowledge access by making research materials available to researchers who might otherwise lack access. This presentation will explore data sharing principles and experiences, highlighting two recent studies: one on smoking behaviors in the United States and another on the association between caregiving and psychological well-being among Florida college students. The resulting datasets were shared under specific terms of use via Harvard Dataverse to facilitate further medical research. In addition, a subset of the tobacco use data and sample code were shared through the Resources Portal of the Teaching of Statistics in the Health Sciences Section of the American Statistical Association to support the teaching of survey methods. The presentation is intended for students, researchers, and practitioners interested in data sharing and open science.
Speakers
avatar for Julia Soulakova

Julia Soulakova

University of Central Florida College of Medicine
Julia Soulakova, Ph.D., is a biostatistician and Professor of Medicine in the Department of Population Health Sciences at the University of Central Florida College of Medicine. Her research interests include statistical methodology with applications to behavioral medicine and social... Read More →
Organizers
avatar for Julia Soulakova

Julia Soulakova

University of Central Florida College of Medicine
Julia Soulakova, Ph.D., is a biostatistician and Professor of Medicine in the Department of Population Health Sciences at the University of Central Florida College of Medicine. Her research interests include statistical methodology with applications to behavioral medicine and social... Read More →
Tuesday October 6, 2026 12:30 - 13:00 UTC
Zoom Room #1

13:00 UTC

S114 - Fractional Statistical Models via Operator Theory: A Data-Driven Framework for Aviation Analytics
Tuesday October 6, 2026 13:00 - 13:30 UTC
Classical statistical models are built upon an assumption of short-range dependence — an assumption that fails dramatically when confronted with the complexity of real-world aviation datasets. Such datasets routinely exhibit long-range memory, non-stationarity, and heavy-tailed distributions that render conventional approaches inadequate. In this work, we propose a novel fractional statistical framework that draws on advanced operator theory to directly address these challenges, offering both rigorous theoretical guarantees and compelling empirical improvements over established baselines.
We construct a family of Toeplitz-type estimators grounded in the theory of α-fractional Bergman spaces, establish their theoretical properties, validate the framework on a large-scale aviation dataset comprising over 500,000 UAE flight records, and demonstrate prediction error reductions of 23–31% over ARIMA and 14–18% over LSTM-based approaches.
Speakers
avatar for Raja'a Alnaimi

Raja'a Alnaimi

emirates aviation university
Dr. Raja’a Al-Naimi is an Assistant Professor in the Department of Mathematics
and Data Science at Emirates Aviation University (EAU), Dubai, UAE. She holds
expertise in operator theory, fractional calculus, and functional analysis, with active
research programs in α-fracti... Read More →
Organizers
avatar for Raja'a Alnaimi

Raja'a Alnaimi

emirates aviation university
Dr. Raja’a Al-Naimi is an Assistant Professor in the Department of Mathematics
and Data Science at Emirates Aviation University (EAU), Dubai, UAE. She holds
expertise in operator theory, fractional calculus, and functional analysis, with active
research programs in α-fracti... Read More →
Tuesday October 6, 2026 13:00 - 13:30 UTC
Zoom Room #1

13:30 UTC

S115 - Reliable Variable Selection for Biomedical Data Science: From Shrinkage Estimation to Interpretable Learning
Tuesday October 6, 2026 13:30 - 14:00 UTC
Modern biomedical data science increasingly relies on datasets with many predictors, limited sample sizes, and complex correlation structures. In these settings, classical regression and standard variable-selection methods may lead to unstable models, overfitting, or conclusions that are difficult to interpret. Shrinkage and penalized estimation provide a principled framework for improving reliability, but their success depends on how multicollinearity and dependence among predictors are handled. This talk discusses reliable variable selection for biomedical data science, moving from classical shrinkage ideas to modern interpretable learning. The presentation will introduce the motivation behind ridge-type methods, LASSO-based procedures, elastic net estimation, and adaptive penalization, with emphasis on approaches that use correlation information to improve model stability and interpretability. Motivated by biomedical applications such as molecular subtype identification and classification problems, the talk will show how statistically grounded regularization can support both prediction and scientific interpretation. The broader message is that reliable biomedical data science requires methods that are not only accurate, but also stable, transparent, reproducible, and interpretable for domain experts.
Speakers
avatar for Mina Norouzirad

Mina Norouzirad

Center for Mathematics and Applications (NOVA Math) and Department of Mathematics, NOVA FCT, Portugal
Mina Norouzirad is an Assistant Researcher in Statistics at the Department of Mathematics and the Center for Mathematics and Applications (NOVA Math), NOVA School of Science and Technology, NOVA University Lisbon, Portugal. She is also Co-Coordinator of the Data Science Thematic Line... Read More →
Organizers
avatar for Mina Norouzirad

Mina Norouzirad

Center for Mathematics and Applications (NOVA Math) and Department of Mathematics, NOVA FCT, Portugal
Mina Norouzirad is an Assistant Researcher in Statistics at the Department of Mathematics and the Center for Mathematics and Applications (NOVA Math), NOVA School of Science and Technology, NOVA University Lisbon, Portugal. She is also Co-Coordinator of the Data Science Thematic Line... Read More →
Tuesday October 6, 2026 13:30 - 14:00 UTC
Zoom Room #1

14:00 UTC

S116 - Data, Equity and Power: Institutional Frameworks for Embedding Young African Women in Decision-Making
Tuesday October 6, 2026 14:00 - 14:30 UTC
Despite advances in data science and statistical training across Africa, a structural disconnect persists between the development of technical talent and its integration into national and regional policy processes. It is particularly pronounced for early-career African women statisticians, who face intersecting institutional barriers,UN Women and PARIS21 shows that women occupy only 15% of chief statistician positions in Sub-Saharan Africa, reflecting patriarchal institutional cultures, limited senior-level mentorship and rigid career progression pathways.

This abstract proposes an actionable policy framework to transition young African women statisticians from technical implementers to strategic policy influencers. Using a comparative case-study methodology through the Young African Statisticians Association network, we examine institutional barriers and entry pathways across selected National Statistical Offices in East and West Africa.

The framework advances three strategic pillars: institutional quotas and fast-track leadership pathways for young women in state-led data initiatives; coordinated mentorship systems linking global statisticians with local networks to institutionalize peer and senior sponsorship and gender-responsive data governance through dedicated advisory positions for young women statisticians in ministerial policy processes.

Embedding young women statisticians in Africa's governance is vital for equitable, evidence-based development.
Speakers
avatar for Sarah Nzioka

Sarah Nzioka

MEL Manager, RefugePoint
A results-oriented MEL Management professional with 8+ years of progressive multi-sector experience in monitoring, evaluation, and managing complex project lifecycles from inception to completion. Possessing proven expertise in designing and implementing tailored M&E & research frameworks... Read More →
Organizers
avatar for Sarah Nzioka

Sarah Nzioka

MEL Manager, RefugePoint
A results-oriented MEL Management professional with 8+ years of progressive multi-sector experience in monitoring, evaluation, and managing complex project lifecycles from inception to completion. Possessing proven expertise in designing and implementing tailored M&E & research frameworks... Read More →
Tuesday October 6, 2026 14:00 - 14:30 UTC
Zoom Room #1

14:30 UTC

S117 - Privacy Doesn't End at the Match: Querying PPRL Data in Practice
Tuesday October 6, 2026 14:30 - 15:00 UTC
Privacy-preserving record linkage (PPRL) has a rich methodological literature on encoding, matching, and cryptographic guarantees — but comparatively little guidance exists on what happens after the match: how do researchers actually access, query, and analyze the linked data that results? This presentation addresses that gap directly, walking through real-world architectures used by statistical agencies, such as the U.S. Census Bureau's linkage key infrastructure and UK Trusted Research Environments, to control researcher access to linked data. We’ll cover how to handle potential uncertainty when analyzing linked data, such as producing confidence tiers, adjusting thresholds, and a novel approach to clerical review within the PPRL system. We’ll also walk through concrete querying structures such as linkage maps and secure enclaves and will explore examples of queries and their outputs. Attendees will leave with a clearer picture of the access-and-analysis landscape for linked data, practical considerations for designing their own linkage projects, and a better sense of how privacy shapes analysis.
Speakers
avatar for Emily Gentles

Emily Gentles

RTI International
Emily Gentles is an expert in data linkage, including entity resolution and privacy-preserving record linkage (PPRL). Ms. Gentles has researched efficient PPRL methods; worked to develop secure linkage systems; and designed innovative record linkage procedures, including manual review... Read More →
Organizers
avatar for Emily Gentles

Emily Gentles

RTI International
Emily Gentles is an expert in data linkage, including entity resolution and privacy-preserving record linkage (PPRL). Ms. Gentles has researched efficient PPRL methods; worked to develop secure linkage systems; and designed innovative record linkage procedures, including manual review... Read More →
Tuesday October 6, 2026 14:30 - 15:00 UTC
Zoom Room #1

18:00 UTC

S312 - New advances on Functional data model-based clustering
Tuesday October 6, 2026 18:00 - 18:30 UTC
Functional data analysis has attracted considerable attention in recent years, and its applications appear in physical processes, genetics, biology, meteorology, and signal processing. Many modern applications produce data best viewed as functions rather than finite-dimensional vectors because of their nature. Beyond the challenges of collecting and preprocessing such data, efficiently handling large volumes of functional observations has become an urgent concern. On one hand, functional data takes values in an infinite-dimensional space, which is challenging to handle with classical methods. On the other hand, ignoring the functionality aspect of data will lead to information loss. In this talk, we highlight model-based clustering methods, a powerful tool in machine learning for identifying subgroup-specific patterns. Specifically, we introduce new model-based clustering techniques for functional data, regardless of the Gaussian assumption. The performance of each algorithm is evaluated through simulations and real-world datasets, and the results confirm their efficiency.
Speakers
avatar for Mina Aminghafari

Mina Aminghafari

Associate Professor, University of Calgary
Dr. Mina Aminghafari is an Associate Professor in the Department of Mathematics and Statistics at the University of Calgary. Her research lies at the intersection of high-dimensional statistics, machine learning, and applied data science, particularly on clustering theory and statistical... Read More →
Organizers
avatar for Mina Aminghafari

Mina Aminghafari

Associate Professor, University of Calgary
Dr. Mina Aminghafari is an Associate Professor in the Department of Mathematics and Statistics at the University of Calgary. Her research lies at the intersection of high-dimensional statistics, machine learning, and applied data science, particularly on clustering theory and statistical... Read More →
Tuesday October 6, 2026 18:00 - 18:30 UTC
Zoom Room #3

18:00 UTC

S408 - Bridging the Gap: How Practicing Data Scientists Use LLMs in the Wild and What It Means for Data Science Education
Tuesday October 6, 2026 18:00 - 18:30 UTC
Since the widespread availability of generative artificial intelligence (GenAI), particularly Large Language Models (LLMs), fundamental questions have emerged about the future of coding in data science. Some predict that data scientists will no longer need traditional coding skills, while others question whether LLMs might replace data scientists entirely. However, these discussions have largely proceeded without empirical evidence of how practicing data scientists actually use these tools.

This study addresses this gap by surveying trained, practicing data scientists to understand if and how they integrate LLMs into their workflows, particularly for writing and editing code and performing other data science tasks. Building on our recent investigation of data science educators' perspectives on LLMs, this research examines real-world usage patterns among practitioners to bridge the gap between current practice and educational preparation.

Our findings will contribute to the data science community in two critical ways. First, by documenting how data scientists are actually working with LLMs four years after their initial release, we provide actionable insights that allow practitioners to learn and adopt effective strategies for integrating these tools into their work. Second, we inform data science education by evaluating whether current pedagogies adequately prepare students for this evolving landscape.
Speakers
avatar for Tiffany Timbers

Tiffany Timbers

University of British Columbia
Dr. Tiffany Timbers is an Associate Professor of Teaching in the Department of Statistics and Instructor in the Master of Data Science program at the University of British Columbia. She holds a PhD in Neuroscience from UBC and completed postdoctoral research in behavioral and neural... Read More →
Organizers
avatar for Tiffany Timbers

Tiffany Timbers

University of British Columbia
Dr. Tiffany Timbers is an Associate Professor of Teaching in the Department of Statistics and Instructor in the Master of Data Science program at the University of British Columbia. She holds a PhD in Neuroscience from UBC and completed postdoctoral research in behavioral and neural... Read More →
Tuesday October 6, 2026 18:00 - 18:30 UTC
Zoom Room #4

18:30 UTC

S409 - Every Step Counts: A journey through crossroads, turning points and learnings
Tuesday October 6, 2026 18:30 - 19:00 UTC
A career is rarely a straight line. It is shaped by choices, unexpected opportunities, challenges, mentors, setbacks, and the willingness to keep learning along the way. In this talk, I reflect on my journey from studying statistics in India to pursuing a PhD in Biostatistics in the United States and building a career as a statistician in the biopharmaceutical industry. Along the way, my experiences have taken me across academic research, internships, clinical development, statistical methodology, multiple therapeutic areas, mentoring, professional service, and leadership within the statistical community.
Rather than focusing only on milestones, this talk explores the crossroads and turning points behind them—the decisions that changed direction, the opportunities that initially seemed small but became important, and the lessons learned from navigating uncertainty and growth. Drawing from experiences in research, clinical trials, interdisciplinary collaboration, mentoring, and professional engagement, I will share how curiosity, adaptability, relationships, and continuous learning have shaped my development as a statistician. The central message is simple: careers are built one step at a time, and even the steps that do not seem significant in the moment can ultimately help define where we go next.

Speakers
avatar for Arinjita Bhattacharyya

Arinjita Bhattacharyya

Arinjita Bhattacharyya, PhD, is a biostatistician and statistical scientist with nearly a decade of experience across the pharmaceutical industry and academia. Most recently an Associate Principal Scientist in Biostatistics at Merck, she has supported clinical development across oncology... Read More →
Tuesday October 6, 2026 18:30 - 19:00 UTC
Zoom Room #4

22:00 UTC

S412 - Building Trustworthy Financial Systems with Explainable AI: Applications in Credit Risk Modeling and Financial Inclusion
Tuesday October 6, 2026 22:00 - Wednesday October 7, 2026 18:30 UTC
Artificial intelligence is transforming how financial institutions assess creditworthiness, detect fraud, and expand access to financial services. Yet many of today's most powerful machine learning models remain opaque, making consequential decisions that customers, regulators, and even financial institutions struggle to interpret. As AI adoption accelerates, ensuring transparency, accountability, and fairness has become essential to building public trust and achieving equitable financial outcomes.
Speakers
avatar for Angela Omogbeme

Angela Omogbeme

University of West Georgia
Angela Omogbeme is a financial technology researcher and data analytics professional specializing in artificial intelligence, explainable AI, fraud detection, credit risk modeling, and financial inclusion.

She holds an M.S. in Business Analytics (4.0 GPA) from the University of West Georgia and an MBA from Edinburgh Business School, Heriot-Watt University, UK. Her research has been presented internationally, including at conferences hosted at the University of Oxford ,UK and the University... Read More →
Organizers
avatar for Angela Omogbeme

Angela Omogbeme

University of West Georgia
Angela Omogbeme is a financial technology researcher and data analytics professional specializing in artificial intelligence, explainable AI, fraud detection, credit risk modeling, and financial inclusion.

She holds an M.S. in Business Analytics (4.0 GPA) from the University of West Georgia and an MBA from Edinburgh Business School, Heriot-Watt University, UK. Her research has been presented internationally, including at conferences hosted at the University of Oxford ,UK and the University... Read More →
Tuesday October 6, 2026 22:00 - Wednesday October 7, 2026 18:30 UTC
Zoom Room #4
 
  • Filter By Date
  • Filter By Venue
  • Filter By Type
  • Country
  • Session ID
  • Timezone


Share Modal

Share this link via

Or copy link

Filter sessions
Apply filters to sessions.
Filtered by Date -