Postdoctoral researcher · Utrecht University

Thales Bertaglia

I study influence and information online: how platforms, creators, and increasingly AI systems shape what reaches people, and how we can study these systems from the outside.

I am a postdoctoral researcher at Utrecht University, where I work on HUMANads, an ERC project on content monetisation and fairness in platform governance. My background is in computer science, but most of my work now sits somewhere between computational social science, platform governance, and law.

Thales Bertaglia

Research

All research

Influence and information

How platforms and creators shape commercial and political information online.

AI-mediated information

How generative AI searches for information, chooses sources, and constructs recommendations.

Studying platforms from the outside

Methods for collecting, auditing, and comparing systems that are only partly observable.

Selected publications

All publications
  1. Preview image for Influencer self-disclosure practices on Instagram: A multi-country longitudinal study

    Influencer self-disclosure practices on Instagram: A multi-country longitudinal study

    Thales Bertaglia, Catalina Goanta, Gerasimos Spanakis, and Adriana Iamnitchi

    Online Social Networks and Media · 2025

    Abstract

    This paper presents a longitudinal study of more than ten years of activity on Instagram consisting of over a million posts by 400 content creators from four countries: the US, Brazil, Netherlands and Germany. Our study shows differences in the professionalisation of content monetisation between countries, yet consistent patterns; significant differences in the frequency of posts yet similar user engagement trends; and significant differences in the disclosure of sponsored content in some countries, with a direct connection with national legislation. We analyse shifts in marketing strategies due to legislative and platform feature changes, focusing on how content creators adapt disclosure methods to different legal environments. We also analyse the impact of disclosures and sponsored posts on engagement and conclude that, although sponsored posts have lower engagement on average, properly disclosing ads does not reduce engagement further. Our observations stress the importance of disclosure compliance and can guide authorities in developing and monitoring them more effectively.

    BibTeX
    @article{bertaglia2025disclosures,
      title = {Influencer self-disclosure practices on Instagram: A multi-country longitudinal study},
      journal = {Online Social Networks and Media},
      volume = {45},
      pages = {100298},
      year = {2025},
      issn = {2468-6964},
      doi = {https://doi.org/10.1016/j.osnem.2024.100298},
      author = {Bertaglia, Thales and Goanta, Catalina and Spanakis, Gerasimos and Iamnitchi, Adriana},
      keywords = {Influencer marketing, Advertising disclosure, Instagram, Self-disclosure practices, Legal compliance},
    }
    
  2. Preview image for InstaSynth: Opportunities and Challenges in Generating Synthetic Instagram Data with ChatGPT for Sponsored Content Detection

    InstaSynth: Opportunities and Challenges in Generating Synthetic Instagram Data with ChatGPT for Sponsored Content Detection

    Thales Bertaglia, Lily Heisig, Rishabh Kaushal, and Adriana Iamnitchi

    Proceedings of the International AAAI Conference on Web and Social Media · 2024

    Abstract

    Large Language Models (LLMs) raise concerns about lowering the cost of generating texts that could be used for unethical or illegal purposes, especially on social media. This paper investigates the promise of such models to help enforce legal requirements related to the disclosure of sponsored content online. We investigate the use of LLMs for generating synthetic Instagram captions with two objectives: The first objective (fidelity) is to produce realistic synthetic datasets. For this, we implement content-level and network-level metrics to assess whether synthetic captions are realistic. The second objective (utility) is to create synthetic data useful for sponsored content detection. For this, we evaluate the effectiveness of the generated synthetic data for training classifiers to identify undisclosed advertisements on Instagram. Our investigations show that the objectives of fidelity and utility may conflict and that prompt engineering is a useful but insufficient strategy. Additionally, we find that while individual synthetic posts may appear realistic, collectively they lack diversity, topic connectivity, and realistic user interaction patterns.

    BibTeX
    @inproceedings{bertaglia2024instasynth,
      title = {InstaSynth: Opportunities and Challenges in Generating Synthetic Instagram Data with ChatGPT for Sponsored Content Detection},
      author = {Bertaglia, Thales and Heisig, Lily and Kaushal, Rishabh and Iamnitchi, Adriana},
      booktitle = {Proceedings of the International AAAI Conference on Web and Social Media},
      volume = {18},
      pages = {139--151},
      doi = {10.1609/icwsm.v18i1.31303},
      year = {2024},
    }
    
  3. Preview image for Closing the Loop: Testing ChatGPT to Generate Model Explanations to Improve Human Labelling of Sponsored Content on Social Media

    Closing the Loop: Testing ChatGPT to Generate Model Explanations to Improve Human Labelling of Sponsored Content on Social Media

    Thales Bertaglia, Stefan Huber, Catalina Goanta, Gerasimos Spanakis, and Adriana Iamnitchi

    Explainable Artificial Intelligence · 2023

    Abstract

    Regulatory bodies worldwide are intensifying their efforts to ensure transparency in influencer marketing on social media through instruments like the Unfair Commercial Practices Directive (UCPD) in the European Union, or Section 5 of the Federal Trade Commission Act. Yet enforcing these obligations has proven to be highly problematic due to the sheer scale of the influencer market. The task of automatically detecting sponsored content aims to enable the monitoring and enforcement of such regulations at scale. Current research in this field primarily frames this problem as a machine learning task, focusing on developing models that achieve high classification performance in detecting ads. These machine learning tasks rely on human data annotation to provide ground truth information. However, agreement between annotators is often low, leading to inconsistent labels that hinder the reliability of models. To improve annotation accuracy and, thus, the detection of sponsored content, we propose using chatGPT to augment the annotation process with phrases identified as relevant features and brief explanations. Our experiments show that this approach consistently improves inter-annotator agreement and annotation accuracy. Additionally, our survey of user experience in the annotation task indicates that the explanations improve the annotators’ confidence and streamline the process. Our proposed methods can ultimately lead to more transparency and alignment with regulatory requirements in sponsored content detection.

    BibTeX
    @inproceedings{bertaglia2023closing,
      author = {Bertaglia, Thales and Huber, Stefan and Goanta, Catalina and Spanakis, Gerasimos and Iamnitchi, Adriana},
      editor = {Longo, Luca},
      title = {Closing the Loop: Testing {ChatGPT} to Generate Model Explanations to Improve Human Labelling of Sponsored Content on Social Media},
      booktitle = {Explainable Artificial Intelligence},
      year = {2023},
      publisher = {Springer Nature Switzerland},
      address = {Cham},
      pages = {198--213},
      isbn = {978-3-031-44067-0},
      doi = {10.1007/978-3-031-44067-0_11},
    }
    
  4. Preview image for Abusive language on social media through the legal looking glass

    Abusive language on social media through the legal looking glass

    Thales Bertaglia, Andreea Grigoriu, Michel Dumontier, and Gijs Dijck

    Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021) · 2021

    Abstract

    Abusive language is a growing phenomenon on social media platforms. Its effects can reach beyond the online context, contributing to mental or emotional stress on users. Automatic tools for detecting abuse can alleviate the issue. In practice, developing automated methods to detect abusive language relies on good quality data. However, there is currently a lack of standards for creating datasets in the field. These standards include definitions of what is considered abusive language, annotation guidelines and reporting on the process. This paper introduces an annotation framework inspired by legal concepts to define abusive language in the context of online harassment. The framework uses a 7-point Likert scale for labelling instead of class labels. We also present ALYT – a dataset of Abusive Language on YouTube. ALYT includes YouTube comments in English extracted from videos on different controversial topics and labelled by Law students. The comments were sampled from the actual collected data, without artificial methods for increasing the abusive content. The paper describes the annotation process thoroughly, including all its guidelines and training steps.

    BibTeX
    @inproceedings{bertaglia2021abusive,
      title = {Abusive language on social media through the legal looking glass},
      author = {Bertaglia, Thales and Grigoriu, Andreea and Dumontier, Michel and van Dijck, Gijs},
      booktitle = {Proceedings of the 5th Workshop on Online Abuse and Harms (WOAH 2021)},
      pages = {191--200},
      year = {2021},
    }
    

Recent updates

All updates
  1. Our paper Towards Fairness Assessment of Dutch Hate Speech Detection has been accepted at WOAH, which will take place at ACL!

  2. Our paper TikTok Search Recommendations: Governance and Research Challenges has been accepted at COMPASS 2025, which will take place at ICWSM!

  3. I defended my PhD on the 7th of November! You can watch the defence here and read my thesis here.

Contact

Email me at contact@thalesbertaglia.com or find my work through the profile links below.