SlideShare a Scribd company logo
1 of 56
Information Retrieval and Social Media
Prof.dr.ir. Arjen P. de Vries
arjen@acm.org
Lecture for the User-Centred Social Media Summer School
Duisburg, September 19, 2017
Social Media
Noun
social media (uncountable)
Interactive forms of media that allow users to interact with and publish to
each other, generally by means of the Internet.
The early 21st century saw a huge increase in social media thanks to the widespread availability of the
Internet.
Social Media
 “Social bookmarking” sites
 “User generated content”
- Images (flickr) and videos (youtube, vimeo), but also blogs, Wikipedia, etc.
 Social network services
- Twitter, facebook, instagram, snapchat
Not just one beast!
User contributed content
Permission based tagging, Set model
Bag model
Global Content
Free for all tagging
Social Media to help improve IR (1)
‘Co-creation’
 Social Media:
- Consumer becomes a co-creator
- Many ‘data consumption’ traces in social media are public
Richer information representations
Richer information representations
 User profiles
- User name, full name, description, image, homepage url, etc.
 Connections between users
- Networks of friends, followers, etc
 Comments/reactions
 Endorsing and sharing
E.g., Twitter
 Bio
- Often includes a geo-location of the profile
 Friends
 Followers
 Lists
- Groups followed Twitter accounts; lists can be followed
 Hashtags
 Mentions
User Demographics
 Gender from Tweet author’s first name
 Geographic location from profile
Diaz, Gamon, Hofman, Kiciman, Rothschild. Online and Social Media as an Imperfect Continuous
Panel Survey. In PLOS ONE, 2016
Detailed User Characteristics…
de Volkskrant, March 13, 2013
Michal Kosinski, David Stillwell, and
Thore Graepel. Private traits and
attributes are predictable from digital
records of human behavior. PNAS
2013.
Youyou, W., Kosinski, M. & Stillwell, D.
(2015) Computer-based personality
judgments are more accurate than
those made by humans. PNAS 2015.
… in Search
 Age and Gender, and perhaps also political and religious
views
 Maps both Page Likes from myPersonality dataset and
search results on a common space of ODP categories
 Learning approach to overcome the difference in
distribution between myPersonality data and Search data
- E.g., their FB dataset has 63% female, vs. only 47% in Bing
Bi, Kosinski, Shokouhi, Graepel. Inferring the Demographics of Search Users. WWW 2013
Many Opportunities for IR
 Expand content representation
 Reduce the vocabulary gap(s) between creators of
content (the indexers) and consumers of content (the
users)
 More diverse views on the same content
LibraryThing
 Items
 People
 Tags
 Ratings
Synonyms
Synonyms
Dissimilar users…
… with similar items
(Pearson Correlation)
Note: this representation ignored the item ratings
Examples
• Humour
• Classic
IR to help improve Social Media
LibraryThing – beyond terms
 Items
 People
 Tags
 Ratings
Maarten Clements, Arjen P. de Vries and Marcel J.T. Reinders. The task
dependent effect of tags and ratings on social media access. TOIS 28, 4, article
21 (November 2010), 42 pages.
Search with Random Walk
 Present nodes according to estimated probability that a
random walk that starts from (task dependent) starting
nodes, would end at this node
Tagging Relationships
Note: this representation used the item ratings in the user – item transitions
An item recommendation walk
Personalized Search
 Assume a user who types a single tag as query
 A soft clustering effect smoothly relates similar concepts
before converging to the background probability
 Homographs like “Java” are disambiguated because the
walk starts in both the query tag and the target user
- So, content that matches the user’s preference is more likely to
be found first
Expert Finding on Twitter
 Empirical evidence demonstrates that a mix of tweet text,
friends, followers and lists is most effective to infer
expertise
 Expertise ground truth taken from Quora, where (many)
users list their expertise and their social media accounts
Xu, Zhou and Lawless. Inferring your expertise from Twitter: combining multiple types of user activity.
WI ‘2017
Multiple Social Networks
 Accounts linked via services like about.me and Quora
 Users explicitly list their multiple accounts in one profile
 Missing data addressed via non-negative matrix
factorization (NMF)
- E.g., 57% list school in FB, 81% in LinkedIn
 Applied to various prediction tasks, e.g.,
topics users are interesting in
Social Media to help improve IR (2)
Relevant for Search… (1/4)
 Wikipedia contains semantically very rich annotations:
- Wikipedia Categories, Lists
- Times (1930, 1931, 1932, etc. etc.)
- Disambiguation pages
- Edit history
Etc.
Note: DBPedia is “just” Wikipedia 
Relevant for Search… (2/4)
 “Twanchor text”
- Tweets citing online media can be used as additional resources
describing the content, just like anchor text
Relevant for Search… (3/4)
 Geotags / POIs
- Recommend geo-locations to people
- Recommend people to geo-locations
- Predict a user’s whereabouts (or “trails”)
Relevant for Search… (4/4)
 Timestamps
- Helps reveal trends, e.g., which documents went viral?
- Allows to search “in the past”
Searching the Social Web
 Do not improve Web search with social annotations, but
improve search in Social
 Builds on the observation in prior work (Goel et al., 2016)
that virality is really different from popularity
- The most viral content is often distinct from the most popular
content being shared online
- Can we surface that content more easily?
Alonso, Kandylas, Tremblay, Hofman, Sen. What’s Happening and What Happened: Searching the
Social Web. WebSci ‘17.
Pipeline
 Content selection:
- Select tweets that contain links and satisfy simple user, content
and time range criteria
 User selection:
- Extract and normalize links and select those that have been
shared by a minimum number of trusted users
 Link selection:
- Clean-up links, compute link virality and popularity, cluster
similar links, and apply heuristic criteria to select good quality
links
 Annotations:
- Generate metadata for the selected links from the associated
tweets
Collecting Data
API Blues
Bit.ly API used in my own research:
/v3/link/content
deprecated
Note: This endpoint was deprecated on 10/15/2014.
API Blues
 The combination of rate limits and Terms of Service of
most social media platforms complicates our life
 Not even to mention volume
- TREC Microblog collection of 2013 “Tweets2013” consists of
107 GB compressed (for only 2 months of data!)
 Did I mention ToS?
- Mandatory continual processing of deletions…
Good News for Twitter
 The Internet Archive distributes two collections from 2013
that can be used as drop-in replacement for evaluation
purposes
 Deletions seem to affect non-relevant documents more
than relevant documents
Sequira and Lin. Finally, a Downloadable Test Collection of Tweets. SIGIR 2017.
Social Media as Panel Survey
 Online population is a non-representative sample of the
off-line world
 Demographic skew and user participation is non-
stationary and difficult to predict over time
- E.g., women are underrepresented in the raw volume of tweets,
but tweet more often about politics than men
- Half of the activity on a specific debate came from individuals
who had not previously posted about the election
Diaz, Gamon, Hofman, Kiciman, Rothschild. Online and Social Media as an Imperfect Continuous
Panel Survey. In PLOS ONE, 2016
Fred Morstatter, Jürgen Pfeffer, Huan Liu and
Kathleen M. Carley. Is the Sample Good
Enough? Comparing Data from Twitter’s
Streaming API with Twitter’s Firehose.
ICWSM 2013
API Blues
Take home message(s)
Take home message(s)
• Social media give access to a rich resource of context
- Including time & location!
Take home message(s)
• Social media give access to a rich resource of context
- Including time & location!
• The academic’s alternative to click data?
Take home message(s)
• Social media give access to a rich resource of context
- Including time & location!
• The academic’s alternative to click data?
• A big open research question:
Can one theory (about matching users and content) address the
complete spectrum of IR tasks that arise in social media?

More Related Content

What's hot

PHOTODYNAMIC THERAPY IN PERIODONTICS.pptx
PHOTODYNAMIC THERAPY IN PERIODONTICS.pptxPHOTODYNAMIC THERAPY IN PERIODONTICS.pptx
PHOTODYNAMIC THERAPY IN PERIODONTICS.pptx
malti19
 
Simplified and modified atraumatic restorative treatment
Simplified and modified atraumatic restorative treatmentSimplified and modified atraumatic restorative treatment
Simplified and modified atraumatic restorative treatment
Hamed Gholami
 

What's hot (20)

Furcation
FurcationFurcation
Furcation
 
Natural Language Processing
Natural Language ProcessingNatural Language Processing
Natural Language Processing
 
Esthetics
EstheticsEsthetics
Esthetics
 
Working length Determination
Working length DeterminationWorking length Determination
Working length Determination
 
School- based oral health education programs; How effective are they?
School- based oral health education programs; How effective are they?School- based oral health education programs; How effective are they?
School- based oral health education programs; How effective are they?
 
HEROIC ENDODONTICS (WHEN TO SAY NO!!)
HEROIC ENDODONTICS (WHEN TO SAY NO!!)HEROIC ENDODONTICS (WHEN TO SAY NO!!)
HEROIC ENDODONTICS (WHEN TO SAY NO!!)
 
SEMINAR ON POST AND CORE pdf.pdf
SEMINAR ON POST AND CORE pdf.pdfSEMINAR ON POST AND CORE pdf.pdf
SEMINAR ON POST AND CORE pdf.pdf
 
PHOTODYNAMIC THERAPY IN PERIODONTICS.pptx
PHOTODYNAMIC THERAPY IN PERIODONTICS.pptxPHOTODYNAMIC THERAPY IN PERIODONTICS.pptx
PHOTODYNAMIC THERAPY IN PERIODONTICS.pptx
 
Natural language processing in artificial intelligence
Natural language processing in artificial intelligenceNatural language processing in artificial intelligence
Natural language processing in artificial intelligence
 
A review and proposed a criteria of success albrektson et al
A review and proposed a criteria of success   albrektson et alA review and proposed a criteria of success   albrektson et al
A review and proposed a criteria of success albrektson et al
 
Minimally invasive endodontics
Minimally invasive endodonticsMinimally invasive endodontics
Minimally invasive endodontics
 
aae_traumaguidelines.pdf
aae_traumaguidelines.pdfaae_traumaguidelines.pdf
aae_traumaguidelines.pdf
 
Platelet rich plasma
Platelet rich plasmaPlatelet rich plasma
Platelet rich plasma
 
Simplified and modified atraumatic restorative treatment
Simplified and modified atraumatic restorative treatmentSimplified and modified atraumatic restorative treatment
Simplified and modified atraumatic restorative treatment
 
Prosthodontics Journal Club: Altered cast impression techniques
Prosthodontics Journal Club: Altered cast impression techniquesProsthodontics Journal Club: Altered cast impression techniques
Prosthodontics Journal Club: Altered cast impression techniques
 
Retreatement finallll/ dental implant courses
Retreatement finallll/ dental implant coursesRetreatement finallll/ dental implant courses
Retreatement finallll/ dental implant courses
 
Implant failure & its management.pptx
Implant failure & its management.pptxImplant failure & its management.pptx
Implant failure & its management.pptx
 
apical regeneration ppt
apical regeneration pptapical regeneration ppt
apical regeneration ppt
 
Fiber-OpticTransillumination In dentistry
Fiber-OpticTransillumination In dentistryFiber-OpticTransillumination In dentistry
Fiber-OpticTransillumination In dentistry
 
Endodontic surgery
Endodontic surgeryEndodontic surgery
Endodontic surgery
 

Similar to Information Retrieval and Social Media

ESSIR 2013 - IR and Social Media
ESSIR 2013 - IR and Social MediaESSIR 2013 - IR and Social Media
ESSIR 2013 - IR and Social Media
Arjen de Vries
 
A review for the online social networks literature
A review for the online social networks literatureA review for the online social networks literature
A review for the online social networks literature
Alexander Decker
 
A review for the online social networks literature
A review for the online social networks literatureA review for the online social networks literature
A review for the online social networks literature
Alexander Decker
 
A research paper on Twitter_Intrinsic versus image related utility in social ...
A research paper on Twitter_Intrinsic versus image related utility in social ...A research paper on Twitter_Intrinsic versus image related utility in social ...
A research paper on Twitter_Intrinsic versus image related utility in social ...
Interskale Digital Marketing and Consulting Pvt. Ltd.
 
The Social Mind Study
The Social Mind StudyThe Social Mind Study
The Social Mind Study
Don Bulmer
 

Similar to Information Retrieval and Social Media (20)

Researching Social Media – Big Data and Social Media Analysis
Researching Social Media – Big Data and Social Media AnalysisResearching Social Media – Big Data and Social Media Analysis
Researching Social Media – Big Data and Social Media Analysis
 
Social media in Research, friend or foe?
Social media in Research, friend or foe?Social media in Research, friend or foe?
Social media in Research, friend or foe?
 
Il laboratorio aperto: limiti e possibilità dell’uso di Facebook, Twitter e Y...
Il laboratorio aperto: limiti e possibilità dell’uso di Facebook, Twitter e Y...Il laboratorio aperto: limiti e possibilità dell’uso di Facebook, Twitter e Y...
Il laboratorio aperto: limiti e possibilità dell’uso di Facebook, Twitter e Y...
 
Utilizing Social Media to Understand People
Utilizing Social Media to Understand PeopleUtilizing Social Media to Understand People
Utilizing Social Media to Understand People
 
WEBINAR: Joining the "buzz": the role of social media in raising research vi...
WEBINAR:  Joining the "buzz": the role of social media in raising research vi...WEBINAR:  Joining the "buzz": the role of social media in raising research vi...
WEBINAR: Joining the "buzz": the role of social media in raising research vi...
 
ESSIR 2013 - IR and Social Media
ESSIR 2013 - IR and Social MediaESSIR 2013 - IR and Social Media
ESSIR 2013 - IR and Social Media
 
Joining the ‘buzz’ : the role of social media in raising research visibility ...
Joining the ‘buzz’ : the role of social media in raising research visibility ...Joining the ‘buzz’ : the role of social media in raising research visibility ...
Joining the ‘buzz’ : the role of social media in raising research visibility ...
 
Helig webinar 6 nov_2014
Helig webinar 6 nov_2014Helig webinar 6 nov_2014
Helig webinar 6 nov_2014
 
A review for the online social networks literature
A review for the online social networks literatureA review for the online social networks literature
A review for the online social networks literature
 
A review for the online social networks literature
A review for the online social networks literatureA review for the online social networks literature
A review for the online social networks literature
 
The Implementation of Social Media for Educational Objectives
The Implementation of Social Media for Educational ObjectivesThe Implementation of Social Media for Educational Objectives
The Implementation of Social Media for Educational Objectives
 
A research paper on Twitter_Intrinsic versus image related utility in social ...
A research paper on Twitter_Intrinsic versus image related utility in social ...A research paper on Twitter_Intrinsic versus image related utility in social ...
A research paper on Twitter_Intrinsic versus image related utility in social ...
 
Research-Open Access-Social Media: a winning combination
Research-Open Access-Social Media: a winning combinationResearch-Open Access-Social Media: a winning combination
Research-Open Access-Social Media: a winning combination
 
Science communication via social media
Science communication via social mediaScience communication via social media
Science communication via social media
 
Science communication via social media
Science communication via social mediaScience communication via social media
Science communication via social media
 
Impact & Interaction: social media as part of communication strategy for rese...
Impact & Interaction: social media as part of communication strategy for rese...Impact & Interaction: social media as part of communication strategy for rese...
Impact & Interaction: social media as part of communication strategy for rese...
 
Vu M Kloos 20071116
Vu M Kloos 20071116Vu M Kloos 20071116
Vu M Kloos 20071116
 
The Social Mind Study
The Social Mind StudyThe Social Mind Study
The Social Mind Study
 
The Social Mind Research Study
The Social Mind Research StudyThe Social Mind Research Study
The Social Mind Research Study
 
Studying Cybercrime: Raising Awareness of Objectivity & Bias
Studying Cybercrime: Raising Awareness of Objectivity & BiasStudying Cybercrime: Raising Awareness of Objectivity & Bias
Studying Cybercrime: Raising Awareness of Objectivity & Bias
 

More from Arjen de Vries

The personal search engine
The personal search engineThe personal search engine
The personal search engine
Arjen de Vries
 
Looking beyond plain text for document representation in the enterprise
Looking beyond plain text for document representation in the enterpriseLooking beyond plain text for document representation in the enterprise
Looking beyond plain text for document representation in the enterprise
Arjen de Vries
 

More from Arjen de Vries (20)

Doing a PhD @ DOSSIER
Doing a PhD @ DOSSIERDoing a PhD @ DOSSIER
Doing a PhD @ DOSSIER
 
Masterclass Big Data (leerlingen)
Masterclass Big Data (leerlingen) Masterclass Big Data (leerlingen)
Masterclass Big Data (leerlingen)
 
Beverwedstrijd Big Data (klas 3/4/5/6)
Beverwedstrijd Big Data (klas 3/4/5/6) Beverwedstrijd Big Data (klas 3/4/5/6)
Beverwedstrijd Big Data (klas 3/4/5/6)
 
Beverwedstrijd Big Data (groep 5/6 en klas 1/2)
Beverwedstrijd Big Data (groep 5/6 en klas 1/2)Beverwedstrijd Big Data (groep 5/6 en klas 1/2)
Beverwedstrijd Big Data (groep 5/6 en klas 1/2)
 
Web Archives and the dream of the Personal Search Engine
Web Archives and the dream of the Personal Search EngineWeb Archives and the dream of the Personal Search Engine
Web Archives and the dream of the Personal Search Engine
 
Information Retrieval intro TMM
Information Retrieval intro TMMInformation Retrieval intro TMM
Information Retrieval intro TMM
 
ACM SIGIR 2017 - Opening - PC Chairs
ACM SIGIR 2017 - Opening - PC ChairsACM SIGIR 2017 - Opening - PC Chairs
ACM SIGIR 2017 - Opening - PC Chairs
 
Data Science Master Specialisation
Data Science Master SpecialisationData Science Master Specialisation
Data Science Master Specialisation
 
PUC Masterclass Big Data
PUC Masterclass Big DataPUC Masterclass Big Data
PUC Masterclass Big Data
 
Bigdata processing with Spark - part II
Bigdata processing with Spark - part IIBigdata processing with Spark - part II
Bigdata processing with Spark - part II
 
Bigdata processing with Spark
Bigdata processing with SparkBigdata processing with Spark
Bigdata processing with Spark
 
TREC 2016: Looking Forward Panel
TREC 2016: Looking Forward PanelTREC 2016: Looking Forward Panel
TREC 2016: Looking Forward Panel
 
The personal search engine
The personal search engineThe personal search engine
The personal search engine
 
Models for Information Retrieval and Recommendation
Models for Information Retrieval and RecommendationModels for Information Retrieval and Recommendation
Models for Information Retrieval and Recommendation
 
Better Contextual Suggestions by Applying Domain Knowledge
Better Contextual Suggestions by Applying Domain KnowledgeBetter Contextual Suggestions by Applying Domain Knowledge
Better Contextual Suggestions by Applying Domain Knowledge
 
Similarity & Recommendation - CWI Scientific Meeting - Sep 27th, 2013
Similarity & Recommendation - CWI Scientific Meeting - Sep 27th, 2013Similarity & Recommendation - CWI Scientific Meeting - Sep 27th, 2013
Similarity & Recommendation - CWI Scientific Meeting - Sep 27th, 2013
 
Looking beyond plain text for document representation in the enterprise
Looking beyond plain text for document representation in the enterpriseLooking beyond plain text for document representation in the enterprise
Looking beyond plain text for document representation in the enterprise
 
Recommendation and Information Retrieval: Two Sides of the Same Coin?
Recommendation and Information Retrieval: Two Sides of the Same Coin?Recommendation and Information Retrieval: Two Sides of the Same Coin?
Recommendation and Information Retrieval: Two Sides of the Same Coin?
 
Searching Political Data by Strategy
Searching Political Data by StrategySearching Political Data by Strategy
Searching Political Data by Strategy
 
How to Search Annotated Text by Strategy?
How to Search Annotated Text by Strategy?How to Search Annotated Text by Strategy?
How to Search Annotated Text by Strategy?
 

Recently uploaded

+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...
+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...
+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...
?#DUbAI#??##{{(☎️+971_581248768%)**%*]'#abortion pills for sale in dubai@
 
development of diagnostic enzyme assay to detect leuser virus
development of diagnostic enzyme assay to detect leuser virusdevelopment of diagnostic enzyme assay to detect leuser virus
development of diagnostic enzyme assay to detect leuser virus
NazaninKarimi6
 
Pests of mustard_Identification_Management_Dr.UPR.pdf
Pests of mustard_Identification_Management_Dr.UPR.pdfPests of mustard_Identification_Management_Dr.UPR.pdf
Pests of mustard_Identification_Management_Dr.UPR.pdf
PirithiRaju
 
Pests of cotton_Borer_Pests_Binomics_Dr.UPR.pdf
Pests of cotton_Borer_Pests_Binomics_Dr.UPR.pdfPests of cotton_Borer_Pests_Binomics_Dr.UPR.pdf
Pests of cotton_Borer_Pests_Binomics_Dr.UPR.pdf
PirithiRaju
 
Digital Dentistry.Digital Dentistryvv.pptx
Digital Dentistry.Digital Dentistryvv.pptxDigital Dentistry.Digital Dentistryvv.pptx
Digital Dentistry.Digital Dentistryvv.pptx
MohamedFarag457087
 

Recently uploaded (20)

9999266834 Call Girls In Noida Sector 22 (Delhi) Call Girl Service
9999266834 Call Girls In Noida Sector 22 (Delhi) Call Girl Service9999266834 Call Girls In Noida Sector 22 (Delhi) Call Girl Service
9999266834 Call Girls In Noida Sector 22 (Delhi) Call Girl Service
 
Forensic Biology & Its biological significance.pdf
Forensic Biology & Its biological significance.pdfForensic Biology & Its biological significance.pdf
Forensic Biology & Its biological significance.pdf
 
+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...
+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...
+971581248768>> SAFE AND ORIGINAL ABORTION PILLS FOR SALE IN DUBAI AND ABUDHA...
 
High Class Escorts in Hyderabad ₹7.5k Pick Up & Drop With Cash Payment 969456...
High Class Escorts in Hyderabad ₹7.5k Pick Up & Drop With Cash Payment 969456...High Class Escorts in Hyderabad ₹7.5k Pick Up & Drop With Cash Payment 969456...
High Class Escorts in Hyderabad ₹7.5k Pick Up & Drop With Cash Payment 969456...
 
COST ESTIMATION FOR A RESEARCH PROJECT.pptx
COST ESTIMATION FOR A RESEARCH PROJECT.pptxCOST ESTIMATION FOR A RESEARCH PROJECT.pptx
COST ESTIMATION FOR A RESEARCH PROJECT.pptx
 
❤Jammu Kashmir Call Girls 8617697112 Personal Whatsapp Number 💦✅.
❤Jammu Kashmir Call Girls 8617697112 Personal Whatsapp Number 💦✅.❤Jammu Kashmir Call Girls 8617697112 Personal Whatsapp Number 💦✅.
❤Jammu Kashmir Call Girls 8617697112 Personal Whatsapp Number 💦✅.
 
Introduction to Viruses
Introduction to VirusesIntroduction to Viruses
Introduction to Viruses
 
Dubai Call Girls Beauty Face Teen O525547819 Call Girls Dubai Young
Dubai Call Girls Beauty Face Teen O525547819 Call Girls Dubai YoungDubai Call Girls Beauty Face Teen O525547819 Call Girls Dubai Young
Dubai Call Girls Beauty Face Teen O525547819 Call Girls Dubai Young
 
module for grade 9 for distance learning
module for grade 9 for distance learningmodule for grade 9 for distance learning
module for grade 9 for distance learning
 
Zoology 5th semester notes( Sumit_yadav).pdf
Zoology 5th semester notes( Sumit_yadav).pdfZoology 5th semester notes( Sumit_yadav).pdf
Zoology 5th semester notes( Sumit_yadav).pdf
 
development of diagnostic enzyme assay to detect leuser virus
development of diagnostic enzyme assay to detect leuser virusdevelopment of diagnostic enzyme assay to detect leuser virus
development of diagnostic enzyme assay to detect leuser virus
 
Pests of mustard_Identification_Management_Dr.UPR.pdf
Pests of mustard_Identification_Management_Dr.UPR.pdfPests of mustard_Identification_Management_Dr.UPR.pdf
Pests of mustard_Identification_Management_Dr.UPR.pdf
 
chemical bonding Essentials of Physical Chemistry2.pdf
chemical bonding Essentials of Physical Chemistry2.pdfchemical bonding Essentials of Physical Chemistry2.pdf
chemical bonding Essentials of Physical Chemistry2.pdf
 
Pests of cotton_Borer_Pests_Binomics_Dr.UPR.pdf
Pests of cotton_Borer_Pests_Binomics_Dr.UPR.pdfPests of cotton_Borer_Pests_Binomics_Dr.UPR.pdf
Pests of cotton_Borer_Pests_Binomics_Dr.UPR.pdf
 
Sector 62, Noida Call girls :8448380779 Model Escorts | 100% verified
Sector 62, Noida Call girls :8448380779 Model Escorts | 100% verifiedSector 62, Noida Call girls :8448380779 Model Escorts | 100% verified
Sector 62, Noida Call girls :8448380779 Model Escorts | 100% verified
 
Site Acceptance Test .
Site Acceptance Test                    .Site Acceptance Test                    .
Site Acceptance Test .
 
Digital Dentistry.Digital Dentistryvv.pptx
Digital Dentistry.Digital Dentistryvv.pptxDigital Dentistry.Digital Dentistryvv.pptx
Digital Dentistry.Digital Dentistryvv.pptx
 
FAIRSpectra - Enabling the FAIRification of Spectroscopy and Spectrometry
FAIRSpectra - Enabling the FAIRification of Spectroscopy and SpectrometryFAIRSpectra - Enabling the FAIRification of Spectroscopy and Spectrometry
FAIRSpectra - Enabling the FAIRification of Spectroscopy and Spectrometry
 
Thyroid Physiology_Dr.E. Muralinath_ Associate Professor
Thyroid Physiology_Dr.E. Muralinath_ Associate ProfessorThyroid Physiology_Dr.E. Muralinath_ Associate Professor
Thyroid Physiology_Dr.E. Muralinath_ Associate Professor
 
High Profile 🔝 8250077686 📞 Call Girls Service in GTB Nagar🍑
High Profile 🔝 8250077686 📞 Call Girls Service in GTB Nagar🍑High Profile 🔝 8250077686 📞 Call Girls Service in GTB Nagar🍑
High Profile 🔝 8250077686 📞 Call Girls Service in GTB Nagar🍑
 

Information Retrieval and Social Media

  • 1. Information Retrieval and Social Media Prof.dr.ir. Arjen P. de Vries arjen@acm.org Lecture for the User-Centred Social Media Summer School Duisburg, September 19, 2017
  • 2. Social Media Noun social media (uncountable) Interactive forms of media that allow users to interact with and publish to each other, generally by means of the Internet. The early 21st century saw a huge increase in social media thanks to the widespread availability of the Internet.
  • 3. Social Media  “Social bookmarking” sites  “User generated content” - Images (flickr) and videos (youtube, vimeo), but also blogs, Wikipedia, etc.  Social network services - Twitter, facebook, instagram, snapchat
  • 4.
  • 5.
  • 6.
  • 7.
  • 8. Not just one beast!
  • 11. Bag model Global Content Free for all tagging
  • 12. Social Media to help improve IR (1)
  • 13. ‘Co-creation’  Social Media: - Consumer becomes a co-creator - Many ‘data consumption’ traces in social media are public
  • 15. Richer information representations  User profiles - User name, full name, description, image, homepage url, etc.  Connections between users - Networks of friends, followers, etc  Comments/reactions  Endorsing and sharing
  • 16. E.g., Twitter  Bio - Often includes a geo-location of the profile  Friends  Followers  Lists - Groups followed Twitter accounts; lists can be followed  Hashtags  Mentions
  • 17. User Demographics  Gender from Tweet author’s first name  Geographic location from profile Diaz, Gamon, Hofman, Kiciman, Rothschild. Online and Social Media as an Imperfect Continuous Panel Survey. In PLOS ONE, 2016
  • 18. Detailed User Characteristics… de Volkskrant, March 13, 2013 Michal Kosinski, David Stillwell, and Thore Graepel. Private traits and attributes are predictable from digital records of human behavior. PNAS 2013. Youyou, W., Kosinski, M. & Stillwell, D. (2015) Computer-based personality judgments are more accurate than those made by humans. PNAS 2015.
  • 19. … in Search  Age and Gender, and perhaps also political and religious views  Maps both Page Likes from myPersonality dataset and search results on a common space of ODP categories  Learning approach to overcome the difference in distribution between myPersonality data and Search data - E.g., their FB dataset has 63% female, vs. only 47% in Bing Bi, Kosinski, Shokouhi, Graepel. Inferring the Demographics of Search Users. WWW 2013
  • 20. Many Opportunities for IR  Expand content representation  Reduce the vocabulary gap(s) between creators of content (the indexers) and consumers of content (the users)  More diverse views on the same content
  • 23. Synonyms Dissimilar users… … with similar items (Pearson Correlation) Note: this representation ignored the item ratings
  • 24.
  • 26. IR to help improve Social Media
  • 27. LibraryThing – beyond terms  Items  People  Tags  Ratings
  • 28. Maarten Clements, Arjen P. de Vries and Marcel J.T. Reinders. The task dependent effect of tags and ratings on social media access. TOIS 28, 4, article 21 (November 2010), 42 pages.
  • 29. Search with Random Walk  Present nodes according to estimated probability that a random walk that starts from (task dependent) starting nodes, would end at this node
  • 31. Note: this representation used the item ratings in the user – item transitions
  • 33. Personalized Search  Assume a user who types a single tag as query
  • 34.  A soft clustering effect smoothly relates similar concepts before converging to the background probability
  • 35.  Homographs like “Java” are disambiguated because the walk starts in both the query tag and the target user - So, content that matches the user’s preference is more likely to be found first
  • 36. Expert Finding on Twitter  Empirical evidence demonstrates that a mix of tweet text, friends, followers and lists is most effective to infer expertise  Expertise ground truth taken from Quora, where (many) users list their expertise and their social media accounts Xu, Zhou and Lawless. Inferring your expertise from Twitter: combining multiple types of user activity. WI ‘2017
  • 37. Multiple Social Networks  Accounts linked via services like about.me and Quora  Users explicitly list their multiple accounts in one profile  Missing data addressed via non-negative matrix factorization (NMF) - E.g., 57% list school in FB, 81% in LinkedIn  Applied to various prediction tasks, e.g., topics users are interesting in
  • 38. Social Media to help improve IR (2)
  • 39. Relevant for Search… (1/4)  Wikipedia contains semantically very rich annotations: - Wikipedia Categories, Lists - Times (1930, 1931, 1932, etc. etc.) - Disambiguation pages - Edit history Etc. Note: DBPedia is “just” Wikipedia 
  • 40. Relevant for Search… (2/4)  “Twanchor text” - Tweets citing online media can be used as additional resources describing the content, just like anchor text
  • 41. Relevant for Search… (3/4)  Geotags / POIs - Recommend geo-locations to people - Recommend people to geo-locations - Predict a user’s whereabouts (or “trails”)
  • 42. Relevant for Search… (4/4)  Timestamps - Helps reveal trends, e.g., which documents went viral? - Allows to search “in the past”
  • 43. Searching the Social Web  Do not improve Web search with social annotations, but improve search in Social  Builds on the observation in prior work (Goel et al., 2016) that virality is really different from popularity - The most viral content is often distinct from the most popular content being shared online - Can we surface that content more easily? Alonso, Kandylas, Tremblay, Hofman, Sen. What’s Happening and What Happened: Searching the Social Web. WebSci ‘17.
  • 44.
  • 45. Pipeline  Content selection: - Select tweets that contain links and satisfy simple user, content and time range criteria  User selection: - Extract and normalize links and select those that have been shared by a minimum number of trusted users  Link selection: - Clean-up links, compute link virality and popularity, cluster similar links, and apply heuristic criteria to select good quality links  Annotations: - Generate metadata for the selected links from the associated tweets
  • 46.
  • 48. API Blues Bit.ly API used in my own research: /v3/link/content deprecated Note: This endpoint was deprecated on 10/15/2014.
  • 49. API Blues  The combination of rate limits and Terms of Service of most social media platforms complicates our life  Not even to mention volume - TREC Microblog collection of 2013 “Tweets2013” consists of 107 GB compressed (for only 2 months of data!)  Did I mention ToS? - Mandatory continual processing of deletions…
  • 50. Good News for Twitter  The Internet Archive distributes two collections from 2013 that can be used as drop-in replacement for evaluation purposes  Deletions seem to affect non-relevant documents more than relevant documents Sequira and Lin. Finally, a Downloadable Test Collection of Tweets. SIGIR 2017.
  • 51. Social Media as Panel Survey  Online population is a non-representative sample of the off-line world  Demographic skew and user participation is non- stationary and difficult to predict over time - E.g., women are underrepresented in the raw volume of tweets, but tweet more often about politics than men - Half of the activity on a specific debate came from individuals who had not previously posted about the election Diaz, Gamon, Hofman, Kiciman, Rothschild. Online and Social Media as an Imperfect Continuous Panel Survey. In PLOS ONE, 2016
  • 52. Fred Morstatter, Jürgen Pfeffer, Huan Liu and Kathleen M. Carley. Is the Sample Good Enough? Comparing Data from Twitter’s Streaming API with Twitter’s Firehose. ICWSM 2013 API Blues
  • 54. Take home message(s) • Social media give access to a rich resource of context - Including time & location!
  • 55. Take home message(s) • Social media give access to a rich resource of context - Including time & location! • The academic’s alternative to click data?
  • 56. Take home message(s) • Social media give access to a rich resource of context - Including time & location! • The academic’s alternative to click data? • A big open research question: Can one theory (about matching users and content) address the complete spectrum of IR tasks that arise in social media?