NLP@ VCU: Identifying adverse effects in English tweets for unbalanced data

Abstract

This paper describes our participation in the Social Media Mining for Health Application (SMM4H 2020) Challenge Track 2 for identifying tweets containing Adverse Effects (AEs). Our system uses Convolutional Neural Networks. We explore downsampling, oversampling, and adjusting the class weights to account for the imbalanced nature of the dataset. Our results showed downsampling outperformed oversampling and adjusting the class weights on the test set however all three obtained similar results on the development set.

Publication
Proceedings of the Fifth Social Media Mining for Health Applications Workshop & Shared Task