Specialisation: Social Media Analysis · Topic Modelling · Sentiment Analysis · Health Communication · Computational Communication
| Affiliation | Department of Communication Sciences, Humanities and International Studies (DISCUI) University of Urbino Carlo Bo, Italy |
| PhD Defended | 22 September 2025 · Cycle XXXVII · Academic Year 2023/2024 |
| Academic Discipline | GSPS-06/A |
| Location | Urbino, Italy 🇮🇹 |
🔍 Open to: Postdoctoral positions · Research fellowships · Visiting researcher roles · Collaborative projects in computational communication, NLP, social media analysis, and digital methods.
I am a computational social scientist and early career researcher whose work sits at the intersection of natural language processing, social media analysis, and computational communication science. I hold a PhD from the University of Urbino Carlo Bo, Italy (Cycle XXXVII, 2025), where I was supervised by Prof. Fabio Giglietto and co-supervised by Prof. Giovanni Boccia Artieri.
My doctoral research examined Facebook Reactions as emotional indicators of public sentiment and engagement with COVID-19 news on Indian media platforms (March 2020 – March 2022), drawing on a dataset of 68,319 Facebook posts from four major English-language outlets. The methodology integrated time-series analysis, BERTopic embedding-based topic modelling, LLM-assisted cluster labelling, and lexicon-based sentiment analysis.
Beyond my dissertation, my research agenda spans health communication, misinformation detection, platform studies, crisis communication, and psychometric scale validation. My published work has received citations tracked live below.
"Facebook Reactions" as Emotional Indicators: A Multi-Method Approach to Analyzing User Engagement with COVID-19 News on Indian Media Platforms Sawood Anwar — University of Urbino Carlo Bo, 2025 🔗 Full text — ORA Uniurb Institutional Repository
The thesis investigates how discrete Facebook Reaction types (Like, Love, Haha, Wow, Sad, Angry) function as affective engagement signals tracking shifting public sentiment across dominant news themes and pandemic phases in India. The dataset includes a focused subset of 8,622 posts covering the early pandemic phase (March 24 – April 14, 2020) alongside the full longitudinal corpus.
Keywords: Facebook Reactions COVID-19 India sentiment analysis BERTopic topic modelling time-series health communication computational communication
[1] Anwar, S., & Giglietto, F. (2024). Facebook Reactions as Emotional Indicators: Analyzing Public Engagement with COVID-19 Pandemic News on Indian Media Platforms During the Early Lockdown Phase. Frontiers in Sociology, 9, 1379265. 🔗 DOI: 10.3389/fsoc.2024.1379265 · 🔓 Open Access
[2] Giglietto, F., Ghasiya, P., Sasahara, K., Anwar, S., & Mincigrucci, R. (2022). Between Localism and Politics: Mapping Coordinated Networks that Circulate Problematic Health Content in India. SSRN Working Paper 4164140. 🔗 DOI: 10.2139/ssrn.4164140 · View on SSRN / Google Scholar
Live counts fetched weekly from Semantic Scholar and Crossref APIs via GitHub Actions. Source data:
citations.json· Updated every Monday at 06:00 UTC
| Source | Live Count | DOI |
|---|---|---|
10.3389/fsoc.2024.1379265 |
||
10.3389/fsoc.2024.1379265 |
||
| View profile → | Full metrics & h-index |
- Computational Communication Science — platform-native engagement signals, affective public response, digital media effects
- Natural Language Processing — topic modelling (BERTopic, STM), text embeddings, LLM-assisted annotation
- Sentiment & Affective Analysis — lexicon-based methods, reaction-type classification, emotion detection
- Health & Crisis Communication — COVID-19, pandemic misinformation, Indian media platforms
- Disinformation & Platform Studies — coordinated inauthentic behaviour, cross-platform analysis
- Quantitative & Computational Methods — time-series, ML classification, survey experiments, psychometric scale validation
| Method Domain | Tools & Packages |
|---|---|
| Topic Modelling | stm, topicmodels, BERTopic (UMAP + HDBSCAN + c-TF-IDF), K-means |
| Sentiment & Affect Analysis | sentimentr, tidytext, AFINN, Bing, NRC lexicons |
| ML & Text Classification | tidymodels, textrecipes, scikit-learn, caret, SHAP |
| Text & Corpus Processing | quanteda, tidyverse, stringr, sentence-transformers |
| LLM Integration | LLM-assisted cluster annotation, OpenAI embeddings |
| Time-Series Analysis | zoo, forecast, anomalize, Z-score, rolling statistics |
| Survey & Psychometrics | psych, lavaan, semTools, srvyr, estimatr, gtsummary |
| Data Collection | CrowdTangle, Meta Content Library API, Reddit API |
| Visualisation & Networks | ggplot2, patchwork, plotly, Gephi |
Open Science Commitment: All code in this repository is publicly available under open licences to support reproducibility and transparency. Primary research data is governed by Meta's CrowdTangle and Content Library data-use agreements.
R Python Facebook COVID-19 India NLP BERTopic Sentiment Analysis Time-Series
Core repository for the doctoral dissertation. Multi-method analysis of Facebook Reactions (68,319 posts, 2020–2022) as emotional indicators of public engagement with COVID-19 news across The Times of India, The Hindu, Indian Express, and Hindustan Times. 📝 Linked publication: Frontiers in Sociology, 2024
R Time-Series Facebook Anomaly Detection Social Media COVID-19
Reproducible R framework for longitudinal social media engagement research: a general-purpose time-series toolkit, a COVID-19 Facebook extension with reaction-type and pandemic-phase stratification, and a health misinformation spike-detection module with event annotation.
R STM Social Media NLP Reproducible Research
Reproducible Structural Topic Model (STM) pipeline for social media corpora. Covers corpus construction, DFM building, searchK model selection, prevalence covariate specification, estimateEffect inference, and publication-quality visualisation.
Python R BERTopic Sentence Transformers UMAP HDBSCAN LLM
BERTopic pipeline for thematic analysis of media corpora: sentence embeddings → UMAP reduction → HDBSCAN clustering → c-TF-IDF topic representations → LLM-based semantic labelling. Outputs exported to R for post-hoc engagement analysis.
R Sentiment Analysis AFINN Bing NRC Lexicon
Systematic comparison of AFINN, Bing, and NRC lexicons applied to social media text, with annotated code and an interpretive guide to lexicon selection for computational communication research.
R Facebook Instagram Health Misinformation STM
Content analysis workflows for studying platform-mediated health communication and misinformation on Meta platforms, combining STM topic modelling with engagement metric analysis.
R Reddit Content Coding Intercoder Reliability Misinformation
Manual content coding framework for political communication and misinformation research on Reddit. Includes codebook, intercoder reliability scripts (Krippendorff's α, Cohen's κ), and descriptive analysis pipeline.
R Facebook Instagram Reddit Cross-Platform
Unified R framework harmonising engagement data from Facebook, Instagram, and Reddit into a single comparative schema for cross-platform digital communication research.
Python R Machine Learning Disinformation SHAP TF-IDF Explainable AI
Supervised ML pipeline benchmarking Logistic Regression, Random Forest, SVM, and Naïve Bayes for disinformation classification in news posts, with SHAP-based model interpretability.
R tidymodels textrecipes News Classification
Supervised text classification pipeline in R for labelling news articles by topic (health, politics, economy) and credibility using tidymodels and textrecipes.
Python R CrowdTangle Meta Content Library API Research Ethics
Documented academic data collection pipeline for Facebook and Instagram research. Covers legacy CrowdTangle exports and Meta Content Library API access, including ethics guidelines and a unified data schema.
R Survey Data Likert Scales srvyr gtsummary Qualtrics
End-to-end survey data analysis pipeline: import, cleaning, Likert scale analysis, descriptive statistics, survey weighting, and diverging bar chart visualisation. Compatible with Qualtrics, SurveyMonkey, SPSS, and Stata exports.
R psych lavaan EFA CFA Psychometrics SEM
Psychometric scale validation using psych and lavaan: item-level inspection, reliability (Cronbach's α, McDonald's ω), EFA with parallel analysis, CFA with fit indices (CFI, TLI, RMSEA, SRMR), and convergent/discriminant validity.
R Survey Experiments Vignette Studies estimatr Causal Inference
Analysis pipeline for survey experiments and vignette studies: randomisation checks, ATE estimation with robust standard errors, OLS regression with modelsummary tables, and moderation/interaction effect visualisation.
| Role | Details |
|---|---|
| Peer Review | Open to reviewing manuscripts in computational communication, NLP, social media analysis, and health communication |
| Open Science | All research code publicly available; committed to reproducible and transparent practices |
| Research Community | Engaged with ICA, AoIR, and computational social science communities |
| Collaboration | Open to cross-disciplinary collaborations in CSS, NLP, platform studies, and global health communication |
| Platform | Link |
|---|---|
| 🌐 Academic Website | sawoodanwar.github.io |
| 🎓 Google Scholar | Sawood Anwar — Google Scholar |
| 🆔 ORCID | 0009-0000-2819-9179 |
| linkedin.com/in/sawood-anwar | |
| 🔗 GitHub | github.com/sawoodanwar |
| anwar1524@gmail.com |
Sawood Anwar · PhD · Computational Social Scientist · University of Urbino Carlo Bo, Italy · 2025