04argmax(project impact)
The problem comes before the model.
Explore six projects by domain. Each case study separates the question, data, method, and reported result—without filling gaps the résumés do not support.
6 projects ∈ portfolio
P01 / Accepted paper · ASONAM 2026
LLM Stance Perception
Women’s Safety Narratives
Measure how stance varies across large-scale social-media discussion of women’s-safety cases in India.
AIMachine LearningData EngineeringResearchData ScienceOpen interactive dashboard ↓
P01 / Accepted paper · ASONAM 2026
Project dashboard
Interactive evidence view
16 cases · 2012–2024 · accuracy 0.7447 · MCC 0.6634
why
The analysis required a defensible bridge from millions of noisy comments to case-level statistical comparisons and a reproducible classification workflow.
data
4.3M+ raw Reddit and YouTube comments; 351,501 comments retained across 16 cases from 2012–2024.
method
Built the processing pipeline, tested distribution differences with Mann–Whitney U, chi-square, and G-tests, then adapted Qwen3.5-9B with LoRA for multi-class stance classification.
toolkit
result
- 0.7447 accuracy
- 0.7450 macro-F1
- 0.6634 MCC
- 3,000-sample held-out evaluation
note
High-confidence error analysis identified systematic model failure modes. The resulting short paper was accepted at ASONAM 2026.
P02 / Nearing completion
BayesV2G
Bayesian Variant-to-Gene Prioritization
Prioritize candidate causal genes for cardiometabolic traits from disease-associated variants.
ResearchData ScienceAnalyticsOpen interactive dashboard ↓
P02 / Nearing completion
Project dashboard
Interactive evidence view
GWAS · cis-eQTL · cis-pQTL · colocalization · cardiometabolic traits
why
Single-evidence and nearest-gene approaches do not express uncertainty across multiple molecular evidence layers.
data
GWAS, cis-eQTL, cis-pQTL, and colocalization evidence for cardiometabolic traits.
method
Estimate posterior variant-to-gene probabilities through uncertainty-aware evidence integration and compare them with nearest-gene and single-evidence baselines.
toolkit
result
- Reproducible statistical-genetics pipeline
- Sensitivity and ablation design
- Baseline comparison framework
note
The project is described as nearing completion; no final performance metric is reported in the résumés.
P03 / Nearing completion
CoreGene-Bayes
Probabilistic Core-Gene Discovery
Identify candidate core disease genes by modeling convergence of genetic perturbations across molecular pathways.
ResearchData ScienceMachine LearningOpen interactive dashboard ↓
P03 / Nearing completion
Project dashboard
Interactive evidence view
Explicitly evaluates conflicting evidence and network-degree bias; no final performance result is claimed.
why
Network evidence can be informative while also carrying uncertainty, conflicting signals, and degree bias.
data
Molecular QTL evidence, Mendelian-randomization evidence, and gene-network information.
method
Estimate posterior core-gene membership and benchmark Bayesian prioritization against centrality and network-propagation approaches.
toolkit
result
- Uncertainty-aware comparison
- Conflicting-evidence analysis
- Network-degree bias evaluation
note
The project is described as nearing completion; no final performance metric is reported in the résumés.
P04 / Methodological development
NeuroMap-GWAS
Spatial Gene-Expression Mapping
Study where neurological disease-associated genes are expressed across human brain regions.
ResearchData ScienceAnalyticsOpen interactive dashboard ↓
P04 / Methodological development
Project dashboard
Interactive evidence view
Includes Parkinson’s, Alzheimer’s, and Huntington’s disease among the 18 conditions.
why
A cross-disease spatial view can test whether genetic susceptibility corresponds to regions affected by disease.
data
GWAS-associated genes and Allen Human Brain Atlas transcriptomic data across 18 neurological and neuropsychiatric diseases.
method
Design regional-enrichment tests with matched random gene sets and empirical null distributions.
toolkit
result
- Standardized 18-disease analysis design
- Brain-region specificity framework
note
The résumés list Parkinson’s, Alzheimer’s, and Huntington’s disease among the 18 conditions.
P05 / Applied AI project
SafeSpell
Harassment & Manipulation Detector
Detect emotionally manipulative and abusive language and present the findings in interpretable form.
AIMachine LearningAnalyticsOpen interactive dashboard ↓
P05 / Applied AI project
Project dashboard
Interactive evidence view
Targets abusive and emotionally manipulative language; no evaluation metric is reported.
why
User-safety review needs structured severity signals instead of an opaque classification alone.
data
Text analyzed for abusive and emotionally manipulative language; the résumés do not name a source dataset.
method
Designed NLP pipelines for keyword detection and severity scoring, then exposed flagged content and trends through a dashboard.
toolkit
result
- Python/FastAPI backend
- Interactive review dashboard
- Structured severity output
note
No date, evaluation metric, demo, or repository link is supplied in the résumés.
P06 / Machine-learning project
MindBill
Emotion-Aware Expense Analytics
Connect small daily purchases with user-labeled emotional outcomes such as regret, happy, and neutral.
Machine LearningAnalyticsData ScienceOpen interactive dashboard ↓
P06 / Machine-learning project
Project dashboard
Interactive evidence view
Regret · happy · neutral · Logistic Regression · Random Forest
why
Expense totals alone do not surface which purchase patterns repeatedly precede regret.
data
Micro-expense records paired with user-labeled emotions.
method
Cleaned and analyzed expense data, trained Logistic Regression and Random Forest models, and designed dashboards for regret trends and triggers.
toolkit
result
- Emotional-outcome prediction workflow
- Regret-trend dashboard
- Spending-trigger analysis
note
No date, evaluation metric, demo, or repository link is supplied in the résumés.