SlideShare a Scribd company logo
1 of 18
Download to read offline
HSE-School of linguistics at Russian Paraphrase
Detection Shared Task
Anastasia Romanova, Mikhail Nefedov
Saint-Petersburg, 2016
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Overview
1 Introduction
2 Task
3 Standard Features
4 Word Embedding Features
5 Results
6 Next steps
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Introduction
Higher School of Economics School of Linguistics
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Task
Compare two sentences
Two types of classification
Standard and Non-standard runs
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Standard Features
Precision
precision =
word-overlap(sentence1,sentence2)
word–count(sentence1)
Recall
recall =
word-overlap(sentence1,sentence2)
word-count(sentence2)
BLEU score
Proposed by IBM (Papineni et al., 2002) for evaluating Machine
Translation Systems
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Standard Features
SyntaxNet
Released by Google in May, 2016
Models for 40 languages
Dependency parse tree
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Standard Features
Tree Edit Distance (Zhang, Shasha, 1989)
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Standard Results
Standard run
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Word Embedding Features
Words as vectors
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Word Embedding Features
Drawbacks of the averaging approach (Rijke
and Kenter, 2015)
Vectors for words Mean vectors
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Word Embedding Features
Before preprocessing
Клинтон выступила с первой речью после поражения на
выборах
After preprocessing
клинтон_S выступать_V первый_A речь_S поражение_S
выбор_S
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Word Embedding Features
BM25 + Word2Vec
sl - longest sentences
ss - shortest sentences
avgsl - average sentence length
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Word Embedding Features
All to all similarities
The boy smiles - The girls laughs
Similarity matrix
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Word Embedding Features
All to all similarities
The boy smiles - The girls laughs
Bins for all values
Bins for maximum values
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Word Embedding Features
Per-dimension similarities
Cosine similarity
Similarity bins
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Results
Non-standard run
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Next steps
Find optimal intervals for bins
Create a new Word2Vec model
Test AdaGram
Compute idf on a larger corpus
Include dependency weighting into BM25
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
Contacts I
anastasiaromane@gmail.com
manefedov26@gmail.com
Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S

More Related Content

Viewers also liked

AINL 2016: Castro, Lopez, Cavalcante, Couto
AINL 2016: Castro, Lopez, Cavalcante, CoutoAINL 2016: Castro, Lopez, Cavalcante, Couto
AINL 2016: Castro, Lopez, Cavalcante, CoutoLidia Pivovarova
 
AINL 2016: Bodrunova, Blekanov, Maksimov
AINL 2016: Bodrunova, Blekanov, MaksimovAINL 2016: Bodrunova, Blekanov, Maksimov
AINL 2016: Bodrunova, Blekanov, MaksimovLidia Pivovarova
 
AINL 2016: Panicheva, Ledovaya
AINL 2016: Panicheva, LedovayaAINL 2016: Panicheva, Ledovaya
AINL 2016: Panicheva, LedovayaLidia Pivovarova
 
AINL 2016: Bastrakova, Ledesma, Millan, Zighed
AINL 2016: Bastrakova, Ledesma, Millan, ZighedAINL 2016: Bastrakova, Ledesma, Millan, Zighed
AINL 2016: Bastrakova, Ledesma, Millan, ZighedLidia Pivovarova
 
AINL 2016: Rykov, Nagornyy, Koltsova, Natta, Kremenets, Manovich, Cerrone, Cr...
AINL 2016: Rykov, Nagornyy, Koltsova, Natta, Kremenets, Manovich, Cerrone, Cr...AINL 2016: Rykov, Nagornyy, Koltsova, Natta, Kremenets, Manovich, Cerrone, Cr...
AINL 2016: Rykov, Nagornyy, Koltsova, Natta, Kremenets, Manovich, Cerrone, Cr...Lidia Pivovarova
 
AINL 2016: Fenogenova, Karpov, Kazorin
AINL 2016: Fenogenova, Karpov, KazorinAINL 2016: Fenogenova, Karpov, Kazorin
AINL 2016: Fenogenova, Karpov, KazorinLidia Pivovarova
 
AINL 2016: Galinsky, Alekseev, Nikolenko
AINL 2016: Galinsky, Alekseev, NikolenkoAINL 2016: Galinsky, Alekseev, Nikolenko
AINL 2016: Galinsky, Alekseev, NikolenkoLidia Pivovarova
 

Viewers also liked (20)

AINL 2016: Castro, Lopez, Cavalcante, Couto
AINL 2016: Castro, Lopez, Cavalcante, CoutoAINL 2016: Castro, Lopez, Cavalcante, Couto
AINL 2016: Castro, Lopez, Cavalcante, Couto
 
AINL 2016: Bodrunova, Blekanov, Maksimov
AINL 2016: Bodrunova, Blekanov, MaksimovAINL 2016: Bodrunova, Blekanov, Maksimov
AINL 2016: Bodrunova, Blekanov, Maksimov
 
AINL 2016: Panicheva, Ledovaya
AINL 2016: Panicheva, LedovayaAINL 2016: Panicheva, Ledovaya
AINL 2016: Panicheva, Ledovaya
 
AINL 2016: Bugaychenko
AINL 2016: BugaychenkoAINL 2016: Bugaychenko
AINL 2016: Bugaychenko
 
AINL 2016: Nikolenko
AINL 2016: NikolenkoAINL 2016: Nikolenko
AINL 2016: Nikolenko
 
AINL 2016: Bastrakova, Ledesma, Millan, Zighed
AINL 2016: Bastrakova, Ledesma, Millan, ZighedAINL 2016: Bastrakova, Ledesma, Millan, Zighed
AINL 2016: Bastrakova, Ledesma, Millan, Zighed
 
AINL 2016: Rykov, Nagornyy, Koltsova, Natta, Kremenets, Manovich, Cerrone, Cr...
AINL 2016: Rykov, Nagornyy, Koltsova, Natta, Kremenets, Manovich, Cerrone, Cr...AINL 2016: Rykov, Nagornyy, Koltsova, Natta, Kremenets, Manovich, Cerrone, Cr...
AINL 2016: Rykov, Nagornyy, Koltsova, Natta, Kremenets, Manovich, Cerrone, Cr...
 
AINL 2016: Ustalov
AINL 2016: Ustalov AINL 2016: Ustalov
AINL 2016: Ustalov
 
AINL 2016: Yagunova
AINL 2016: YagunovaAINL 2016: Yagunova
AINL 2016: Yagunova
 
AINL 2016: Eyecioglu
AINL 2016: EyeciogluAINL 2016: Eyecioglu
AINL 2016: Eyecioglu
 
AINL 2016: Fenogenova, Karpov, Kazorin
AINL 2016: Fenogenova, Karpov, KazorinAINL 2016: Fenogenova, Karpov, Kazorin
AINL 2016: Fenogenova, Karpov, Kazorin
 
AINL 2016: Galinsky, Alekseev, Nikolenko
AINL 2016: Galinsky, Alekseev, NikolenkoAINL 2016: Galinsky, Alekseev, Nikolenko
AINL 2016: Galinsky, Alekseev, Nikolenko
 
AINL 2016: Goncharov
AINL 2016: GoncharovAINL 2016: Goncharov
AINL 2016: Goncharov
 
AINL 2016: Kravchenko
AINL 2016: KravchenkoAINL 2016: Kravchenko
AINL 2016: Kravchenko
 
AINL 2016: Moskvichev
AINL 2016: MoskvichevAINL 2016: Moskvichev
AINL 2016: Moskvichev
 
AINL 2016: Strijov
AINL 2016: StrijovAINL 2016: Strijov
AINL 2016: Strijov
 
AINL 2016: Maraev
AINL 2016: MaraevAINL 2016: Maraev
AINL 2016: Maraev
 
AINL 2016: Khudobakhshov
AINL 2016: KhudobakhshovAINL 2016: Khudobakhshov
AINL 2016: Khudobakhshov
 
AINL 2016: Malykh
AINL 2016: MalykhAINL 2016: Malykh
AINL 2016: Malykh
 
AINL 2016: Filchenkov
AINL 2016: FilchenkovAINL 2016: Filchenkov
AINL 2016: Filchenkov
 

More from Lidia Pivovarova

Classification and clustering in media monitoring: from knowledge engineering...
Classification and clustering in media monitoring: from knowledge engineering...Classification and clustering in media monitoring: from knowledge engineering...
Classification and clustering in media monitoring: from knowledge engineering...Lidia Pivovarova
 
Convolutional neural networks for text classification
Convolutional neural networks for text classificationConvolutional neural networks for text classification
Convolutional neural networks for text classificationLidia Pivovarova
 
Grouping business news stories based on salience of named entities
Grouping business news stories based on salience of named entitiesGrouping business news stories based on salience of named entities
Grouping business news stories based on salience of named entitiesLidia Pivovarova
 
Интеллектуальный анализ текста
Интеллектуальный анализ текстаИнтеллектуальный анализ текста
Интеллектуальный анализ текстаLidia Pivovarova
 
AINL 2016: Shavrina, Selegey
AINL 2016: Shavrina, SelegeyAINL 2016: Shavrina, Selegey
AINL 2016: Shavrina, SelegeyLidia Pivovarova
 

More from Lidia Pivovarova (8)

Classification and clustering in media monitoring: from knowledge engineering...
Classification and clustering in media monitoring: from knowledge engineering...Classification and clustering in media monitoring: from knowledge engineering...
Classification and clustering in media monitoring: from knowledge engineering...
 
Convolutional neural networks for text classification
Convolutional neural networks for text classificationConvolutional neural networks for text classification
Convolutional neural networks for text classification
 
Grouping business news stories based on salience of named entities
Grouping business news stories based on salience of named entitiesGrouping business news stories based on salience of named entities
Grouping business news stories based on salience of named entities
 
Интеллектуальный анализ текста
Интеллектуальный анализ текстаИнтеллектуальный анализ текста
Интеллектуальный анализ текста
 
AINL 2016: Shavrina, Selegey
AINL 2016: Shavrina, SelegeyAINL 2016: Shavrina, Selegey
AINL 2016: Shavrina, Selegey
 
AINL 2016:
AINL 2016: AINL 2016:
AINL 2016:
 
AINL 2016: Grigorieva
AINL 2016: GrigorievaAINL 2016: Grigorieva
AINL 2016: Grigorieva
 
AINL 2016: Just AI
AINL 2016: Just AIAINL 2016: Just AI
AINL 2016: Just AI
 

Recently uploaded

Disentangling the origin of chemical differences using GHOST
Disentangling the origin of chemical differences using GHOSTDisentangling the origin of chemical differences using GHOST
Disentangling the origin of chemical differences using GHOSTSérgio Sacani
 
Module 4: Mendelian Genetics and Punnett Square
Module 4:  Mendelian Genetics and Punnett SquareModule 4:  Mendelian Genetics and Punnett Square
Module 4: Mendelian Genetics and Punnett SquareIsiahStephanRadaza
 
Recombinant DNA technology( Transgenic plant and animal)
Recombinant DNA technology( Transgenic plant and animal)Recombinant DNA technology( Transgenic plant and animal)
Recombinant DNA technology( Transgenic plant and animal)DHURKADEVIBASKAR
 
GFP in rDNA Technology (Biotechnology).pptx
GFP in rDNA Technology (Biotechnology).pptxGFP in rDNA Technology (Biotechnology).pptx
GFP in rDNA Technology (Biotechnology).pptxAleenaTreesaSaji
 
Boyles law module in the grade 10 science
Boyles law module in the grade 10 scienceBoyles law module in the grade 10 science
Boyles law module in the grade 10 sciencefloriejanemacaya1
 
Luciferase in rDNA technology (biotechnology).pptx
Luciferase in rDNA technology (biotechnology).pptxLuciferase in rDNA technology (biotechnology).pptx
Luciferase in rDNA technology (biotechnology).pptxAleenaTreesaSaji
 
zoogeography of pakistan.pptx fauna of Pakistan
zoogeography of pakistan.pptx fauna of Pakistanzoogeography of pakistan.pptx fauna of Pakistan
zoogeography of pakistan.pptx fauna of Pakistanzohaibmir069
 
Analytical Profile of Coleus Forskohlii | Forskolin .pdf
Analytical Profile of Coleus Forskohlii | Forskolin .pdfAnalytical Profile of Coleus Forskohlii | Forskolin .pdf
Analytical Profile of Coleus Forskohlii | Forskolin .pdfSwapnil Therkar
 
Animal Communication- Auditory and Visual.pptx
Animal Communication- Auditory and Visual.pptxAnimal Communication- Auditory and Visual.pptx
Animal Communication- Auditory and Visual.pptxUmerFayaz5
 
A relative description on Sonoporation.pdf
A relative description on Sonoporation.pdfA relative description on Sonoporation.pdf
A relative description on Sonoporation.pdfnehabiju2046
 
STERILITY TESTING OF PHARMACEUTICALS ppt by DR.C.P.PRINCE
STERILITY TESTING OF PHARMACEUTICALS ppt by DR.C.P.PRINCESTERILITY TESTING OF PHARMACEUTICALS ppt by DR.C.P.PRINCE
STERILITY TESTING OF PHARMACEUTICALS ppt by DR.C.P.PRINCEPRINCE C P
 
Discovery of an Accretion Streamer and a Slow Wide-angle Outflow around FUOri...
Discovery of an Accretion Streamer and a Slow Wide-angle Outflow around FUOri...Discovery of an Accretion Streamer and a Slow Wide-angle Outflow around FUOri...
Discovery of an Accretion Streamer and a Slow Wide-angle Outflow around FUOri...Sérgio Sacani
 
Work, Energy and Power for class 10 ICSE Physics
Work, Energy and Power for class 10 ICSE PhysicsWork, Energy and Power for class 10 ICSE Physics
Work, Energy and Power for class 10 ICSE Physicsvishikhakeshava1
 
Neurodevelopmental disorders according to the dsm 5 tr
Neurodevelopmental disorders according to the dsm 5 trNeurodevelopmental disorders according to the dsm 5 tr
Neurodevelopmental disorders according to the dsm 5 trssuser06f238
 
Call Girls in Munirka Delhi 💯Call Us 🔝8264348440🔝
Call Girls in Munirka Delhi 💯Call Us 🔝8264348440🔝Call Girls in Munirka Delhi 💯Call Us 🔝8264348440🔝
Call Girls in Munirka Delhi 💯Call Us 🔝8264348440🔝soniya singh
 
Call Girls in Mayapuri Delhi 💯Call Us 🔝9953322196🔝 💯Escort.
Call Girls in Mayapuri Delhi 💯Call Us 🔝9953322196🔝 💯Escort.Call Girls in Mayapuri Delhi 💯Call Us 🔝9953322196🔝 💯Escort.
Call Girls in Mayapuri Delhi 💯Call Us 🔝9953322196🔝 💯Escort.aasikanpl
 
PossibleEoarcheanRecordsoftheGeomagneticFieldPreservedintheIsuaSupracrustalBe...
PossibleEoarcheanRecordsoftheGeomagneticFieldPreservedintheIsuaSupracrustalBe...PossibleEoarcheanRecordsoftheGeomagneticFieldPreservedintheIsuaSupracrustalBe...
PossibleEoarcheanRecordsoftheGeomagneticFieldPreservedintheIsuaSupracrustalBe...Sérgio Sacani
 

Recently uploaded (20)

The Philosophy of Science
The Philosophy of ScienceThe Philosophy of Science
The Philosophy of Science
 
9953056974 Young Call Girls In Mahavir enclave Indian Quality Escort service
9953056974 Young Call Girls In Mahavir enclave Indian Quality Escort service9953056974 Young Call Girls In Mahavir enclave Indian Quality Escort service
9953056974 Young Call Girls In Mahavir enclave Indian Quality Escort service
 
Disentangling the origin of chemical differences using GHOST
Disentangling the origin of chemical differences using GHOSTDisentangling the origin of chemical differences using GHOST
Disentangling the origin of chemical differences using GHOST
 
Module 4: Mendelian Genetics and Punnett Square
Module 4:  Mendelian Genetics and Punnett SquareModule 4:  Mendelian Genetics and Punnett Square
Module 4: Mendelian Genetics and Punnett Square
 
Recombinant DNA technology( Transgenic plant and animal)
Recombinant DNA technology( Transgenic plant and animal)Recombinant DNA technology( Transgenic plant and animal)
Recombinant DNA technology( Transgenic plant and animal)
 
Engler and Prantl system of classification in plant taxonomy
Engler and Prantl system of classification in plant taxonomyEngler and Prantl system of classification in plant taxonomy
Engler and Prantl system of classification in plant taxonomy
 
GFP in rDNA Technology (Biotechnology).pptx
GFP in rDNA Technology (Biotechnology).pptxGFP in rDNA Technology (Biotechnology).pptx
GFP in rDNA Technology (Biotechnology).pptx
 
Boyles law module in the grade 10 science
Boyles law module in the grade 10 scienceBoyles law module in the grade 10 science
Boyles law module in the grade 10 science
 
Luciferase in rDNA technology (biotechnology).pptx
Luciferase in rDNA technology (biotechnology).pptxLuciferase in rDNA technology (biotechnology).pptx
Luciferase in rDNA technology (biotechnology).pptx
 
zoogeography of pakistan.pptx fauna of Pakistan
zoogeography of pakistan.pptx fauna of Pakistanzoogeography of pakistan.pptx fauna of Pakistan
zoogeography of pakistan.pptx fauna of Pakistan
 
Analytical Profile of Coleus Forskohlii | Forskolin .pdf
Analytical Profile of Coleus Forskohlii | Forskolin .pdfAnalytical Profile of Coleus Forskohlii | Forskolin .pdf
Analytical Profile of Coleus Forskohlii | Forskolin .pdf
 
Animal Communication- Auditory and Visual.pptx
Animal Communication- Auditory and Visual.pptxAnimal Communication- Auditory and Visual.pptx
Animal Communication- Auditory and Visual.pptx
 
A relative description on Sonoporation.pdf
A relative description on Sonoporation.pdfA relative description on Sonoporation.pdf
A relative description on Sonoporation.pdf
 
STERILITY TESTING OF PHARMACEUTICALS ppt by DR.C.P.PRINCE
STERILITY TESTING OF PHARMACEUTICALS ppt by DR.C.P.PRINCESTERILITY TESTING OF PHARMACEUTICALS ppt by DR.C.P.PRINCE
STERILITY TESTING OF PHARMACEUTICALS ppt by DR.C.P.PRINCE
 
Discovery of an Accretion Streamer and a Slow Wide-angle Outflow around FUOri...
Discovery of an Accretion Streamer and a Slow Wide-angle Outflow around FUOri...Discovery of an Accretion Streamer and a Slow Wide-angle Outflow around FUOri...
Discovery of an Accretion Streamer and a Slow Wide-angle Outflow around FUOri...
 
Work, Energy and Power for class 10 ICSE Physics
Work, Energy and Power for class 10 ICSE PhysicsWork, Energy and Power for class 10 ICSE Physics
Work, Energy and Power for class 10 ICSE Physics
 
Neurodevelopmental disorders according to the dsm 5 tr
Neurodevelopmental disorders according to the dsm 5 trNeurodevelopmental disorders according to the dsm 5 tr
Neurodevelopmental disorders according to the dsm 5 tr
 
Call Girls in Munirka Delhi 💯Call Us 🔝8264348440🔝
Call Girls in Munirka Delhi 💯Call Us 🔝8264348440🔝Call Girls in Munirka Delhi 💯Call Us 🔝8264348440🔝
Call Girls in Munirka Delhi 💯Call Us 🔝8264348440🔝
 
Call Girls in Mayapuri Delhi 💯Call Us 🔝9953322196🔝 💯Escort.
Call Girls in Mayapuri Delhi 💯Call Us 🔝9953322196🔝 💯Escort.Call Girls in Mayapuri Delhi 💯Call Us 🔝9953322196🔝 💯Escort.
Call Girls in Mayapuri Delhi 💯Call Us 🔝9953322196🔝 💯Escort.
 
PossibleEoarcheanRecordsoftheGeomagneticFieldPreservedintheIsuaSupracrustalBe...
PossibleEoarcheanRecordsoftheGeomagneticFieldPreservedintheIsuaSupracrustalBe...PossibleEoarcheanRecordsoftheGeomagneticFieldPreservedintheIsuaSupracrustalBe...
PossibleEoarcheanRecordsoftheGeomagneticFieldPreservedintheIsuaSupracrustalBe...
 

AINL 2016: Romanova, Nefedov

  • 1. HSE-School of linguistics at Russian Paraphrase Detection Shared Task Anastasia Romanova, Mikhail Nefedov Saint-Petersburg, 2016 Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 2. Overview 1 Introduction 2 Task 3 Standard Features 4 Word Embedding Features 5 Results 6 Next steps Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 3. Introduction Higher School of Economics School of Linguistics Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 4. Task Compare two sentences Two types of classification Standard and Non-standard runs Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 5. Standard Features Precision precision = word-overlap(sentence1,sentence2) word–count(sentence1) Recall recall = word-overlap(sentence1,sentence2) word-count(sentence2) BLEU score Proposed by IBM (Papineni et al., 2002) for evaluating Machine Translation Systems Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 6. Standard Features SyntaxNet Released by Google in May, 2016 Models for 40 languages Dependency parse tree Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 7. Standard Features Tree Edit Distance (Zhang, Shasha, 1989) Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 8. Standard Results Standard run Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 9. Word Embedding Features Words as vectors Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 10. Word Embedding Features Drawbacks of the averaging approach (Rijke and Kenter, 2015) Vectors for words Mean vectors Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 11. Word Embedding Features Before preprocessing Клинтон выступила с первой речью после поражения на выборах After preprocessing клинтон_S выступать_V первый_A речь_S поражение_S выбор_S Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 12. Word Embedding Features BM25 + Word2Vec sl - longest sentences ss - shortest sentences avgsl - average sentence length Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 13. Word Embedding Features All to all similarities The boy smiles - The girls laughs Similarity matrix Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 14. Word Embedding Features All to all similarities The boy smiles - The girls laughs Bins for all values Bins for maximum values Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 15. Word Embedding Features Per-dimension similarities Cosine similarity Similarity bins Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 16. Results Non-standard run Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 17. Next steps Find optimal intervals for bins Create a new Word2Vec model Test AdaGram Compute idf on a larger corpus Include dependency weighting into BM25 Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S
  • 18. Contacts I anastasiaromane@gmail.com manefedov26@gmail.com Anastasia Romanova, Mikhail Nefedov HSE-School of linguistics at Russian Paraphrase Detection S