Abstract

Social media has become the first source of information for many people. The amount of information posted on social media daily has become very vast that it became difficult to track. One of the most popular social media applications is Twitter. Users follow lots of news accounts, public figures, and their friends so they can be updated by the latest events around them. Since the dialect language and the style of writing differ from a region to another, our objective in this research is to extract trending topics for an Egyptian twitter user. In this way, the user can easily get at a glimpse of the trending topics discussed by the people he follows. To find the best approach achieving our objective, we investigate the document pivot and the feature pivot approaches. By applying the document pivot approach on the baseline data using tf-itf (term frequency-inverse tweet frequency) representation, repeated bisecting k-means clustering technique and extracting most frequent n-grams from each cluster we could achieve a recall value of 100% and F1 measure of 0.8. The application of the feature pivot approach on the baseline data using the content similarity algorithm to group related unigrams together, could achieve a recall value of 100% and F1 measure of 0.923. To validate our results we collected 12 different data sets of different sizes (200, 400, 600, and 1200) and from three different domains (sports, entertainment, and news) then applied both approaches to them. The average recall, precision and F1 measure values resulted from applying the feature pivot approach are larger than those achieved by applying the document pivot approach. To make sure this difference in results is statistically significant we applied the Two-sample one-tailed paired significance t-test that showed the results are significantly better at confidence interval of 90% The results showed that the document pivot approach could extract the trending topics for an Egyptian twitter user with an average recall value of 0.714, average precision value of 0.521, and average F1 measure value of 0.556 versus average recall, precision and F1 measure values of 0.981, 0.754, and 0.833 respectively, when applying the feature pivot approach. â€ƒ

Department

Computer Science & Engineering Department

Degree Name

MS in Computer Science

Graduation Date

6-1-2016

Submission Date

April 2016

First Advisor

Rafea, Ahmed

Committee Member 1

Aly, Sherif G.

Committee Member 2

El Kadi, Amr

Extent

103 p.

Document Type

Master's Thesis

Rights

The author retains all rights with regard to copyright. The author certifies that written permission from the owner(s) of third-party copyrighted matter included in the thesis, dissertation, paper, or record of study has been obtained. The author further certifies that IRB approval has been obtained for this thesis, or that IRB approval is not necessary for this thesis. Insofar as this thesis, dissertation, paper, or record of study is an educational record as defined in the Family Educational Rights and Privacy Act (FERPA) (20 USC 1232g), the author has granted consent to disclosure of it to anyone who requests a copy.

Institutional Review Board (IRB) Approval

Approval has been obtained for this item

Recommended Citation

APA Citation

Mostafa, N. (2016).Trending topic extraction from social media [Master's Thesis, the American University in Cairo]. AUC Knowledge Fountain.
https://fount.aucegypt.edu/etds/246

MLA Citation

Mostafa, Nada Ayman A.. Trending topic extraction from social media. 2016. American University in Cairo, Master's Thesis. AUC Knowledge Fountain.
https://fount.aucegypt.edu/etds/246

Download

COinS

Theses and Dissertations

Trending topic extraction from social media

Abstract

Department

Degree Name

Graduation Date

Submission Date

First Advisor

Committee Member 1

Committee Member 2

Extent

Document Type

Rights

Institutional Review Board (IRB) Approval

Recommended Citation

APA Citation

MLA Citation

Search

Browse

Submit

Theses and Dissertations

Trending topic extraction from social media

Author

Abstract

Department

Degree Name

Graduation Date

Submission Date

First Advisor

Committee Member 1

Committee Member 2

Extent

Document Type

Rights

Institutional Review Board (IRB) Approval

Recommended Citation

APA Citation

MLA Citation

Share

Search

Browse

Submit