Showing posts with label IFI7167. Show all posts
Showing posts with label IFI7167. Show all posts

Wednesday, May 7, 2014

Social Computing Data Analysis - Will Facebook die out by 2017

The paper titled Epidemiological modeling of online social network dynamics by John Cannarella & Joshua A. Spechler (2014) uses epidemic recovery models (irSIR - infection recovery S = number susceptible, I =number infectious, R =number recovered) to predict the rise and fall of Online Social Networks using Google Trends search data in place of actual usage reports. It first tries to fit the model with the rise and fall of MySpace usage, then uses the adapted model to predict if the same effects apply to Facebook data and, therefore, when Facebook would see similar decline.



The usage of search terms to predict the spread of disease has been proven before and is the basis for Google's Flu Trends, which attempts to predict when and where flu epidemics will hit next. But the model does not exactly fit OSN adoption and the paper discusses some of the shortcoming of the model, but not all. The first shortcoming discussed by the paper is that, differently from diseases, people do not join an OSN expecting and/or making a conscious effort to leave them. People will remain members for as long as it is interesting to them, meaning the as long as there are enough friends using an OSN to justify being a member, they'll stay.


Another shortcoming (one that is not discussed satisfactorily in the paper) is that R, the number of people who recovered from a disease, is replaced by people opposed to join an OSN, both those who never joined in the first place and those who joined and left deciding to never come back. This adaptation is necessary because it assumes S+I+R = N, meaning it assumes population remains constant during the study and adding up susceptible, infected and recovered members gives you the full population. This carries two shortcomings:

First, you cannot consciously decide to resist/accept to be infected by a disease, but people make conscious decisions about joining/not joining an OSN all the time, decisions that can change with time (natural immunity can fluctuate, but not be switched on and off at will);
Second, internet usage numbers in the period have not remained constant, instead growing exponentially. The assumption of S+I+R = N is feeble at best.



The third (or fourth?) shortcoming is that the data includes a highly eschewed input, as Google Trends data shows a circa 20% jump around October 2012 that never recovered. This data was 'corrected' by multiplying all input after that date by a correction factor derived from their own projections of where the data points should be, without feedback from Google on what exactly is the nature of the change in data. The turning point in the search data occurs only after the correction factor is applied, putting into question how much of the reduction observed is actually bias generated by the researchers' own 'correction' of input data.

All in all, the paper interestingly draws a parallel between the decline of MySpace and historical Facebook usage data, but the predictions derived by the parallel must be taken with a big grain of salt.

Reference

John Cannarella & Joshua A. Spechler, Epidemiological modeling of online social network dynamics (2014)

Images: John Cannarella & Joshua A. Spechler, Epidemiological modeling of online social network dynamics (2014)

Pingback

Wednesday, February 26, 2014

Operation War Diary - Case Exercise

Crowdsourcing, a term coined by Jeff Howe, author of the seminal Wired article "The Rise of Crowdsourcing", was an attempt to explain the trend among internet companies in the first half of the 2000 decade of tapping into their own public for content. in a 2008 article for Convergence, Daren Brabhan of the University of Utah defined it as "an online, distributed problem-solving and production model".

Many different definitions exist, but in a nutshell, any initiative, network or service where content or other facet of the product(s) is/are created by the general public can be defined as a type of crowdsourcing. Many of these include gaming components (badges, scores, etc) to generate incentives to users.



Some examples:

Foursquare
A location-based social networking where users "check-in" to venues. The gamified interface drives users to input information on their favorite places, generating a comprehensive database of the best places in town that can be filtered by interest.

Reddit
A news and entertainment website where users generate the content, comment and also vote news stories up and down, effectively replacing all parts of the editorial team.

Quirky
A website where users create new products that the company later manufactures and sells. Users suggest products, vote for features and even choose the name, price and tag lines for the ones that finally go on sale.

One of the most accessed websites in the world, Wikipedia is an encyclopedia created and maintained by it's own users, who edit articles, verify sources and self-regulate the behavior of their own colleagues.

Operation War Diary brings this approach to cataloging and categorizing the diaries of several British military units in the First World War. These diaries are part of the National Archives catalog and are being digitized as part of the 100th anniversary of the Great War.

Users get access to these digitized diaries and, in turn, tag the pages indicating what parts of the pages refer to which days, every time a person is mentioned (first name, last name, rank and reason for being mentioned), any time a given place is mentioned, to what category a given entry belongs etc. This information will enable the researchers at the National Archives to better index and analyse the content of all the diaries being digitized.

In Wikipedia, where user-generated content is directly available to other users and affects the experience. For this reason, there's need for moderation (reviewing, content control and reversion of modifications) to control the quality of the generated content. In Operation War Diary, the quality of the content is obtained through redundancy: as many different users will tag the same pages, the level of confidence on any given piece of content can be measured by how many users agree or disagree on it. This way, the more users collaborate, the more trustworthy the final content is.

References

Brabham, D.C. Crowdsourcing as a model for problem solving. Convergence: The International Journal of Research into New Media Technologies, 14(1), February 2008.

Howe, Jeff. The Rise of Crowdsourcing. Wired., 2006.

Deterding, S. et alli. "From game design elements to gamefulness: Defining "gamification"". Proceedings of the 15th International Academic MindTrek Conference. pp. 9–15, 2011

Oomen, J., and Aroyo, L. "Crowdsourcing in the cultural heritage domain: opportunities and challenges." In Proceedings of the 5th International Conference on Communities and Technologies, pp. 138-149. ACM, 2011.


Pingback