Skip to main content
arXiv is now an independent nonprofit! Learn more

Showing 1–17 of 17 results for author: Rama, T

Searching in archive cs. Search in all archives.
.
  1. arXiv:2406.14680  [pdf, other

    cs.CL

    Dravidian language family through Universal Dependencies lens

    Authors: Taraka Rama, Sowmya Vajjala

    Abstract: The Universal Dependencies (UD) project aims to create a cross-linguistically consistent dependency annotation for multiple languages, to facilitate multilingual NLP. It currently supports 114 languages. Dravidian languages are spoken by over 200 million people across the word, and yet there are only two languages from this family in UD. This paper examines some of the morphological and syntactic… ▽ More

    Submitted 20 June, 2024; originally announced June 2024.

    Comments: unpublished report from 2021

  2. arXiv:2402.02807  [pdf, other

    cs.CL cs.SD eess.AS

    Are Sounds Sound for Phylogenetic Reconstruction?

    Authors: Luise Häuser, Gerhard Jäger, Taraka Rama, Johann-Mattis List, Alexandros Stamatakis

    Abstract: In traditional studies on language evolution, scholars often emphasize the importance of sound laws and sound correspondences for phylogenetic inference of language family trees. However, to date, computational approaches have typically not taken this potential into account. Most computational studies still rely on lexical cognates as major data source for phylogenetic reconstruction in linguistic… ▽ More

    Submitted 14 May, 2024; v1 submitted 5 February, 2024; originally announced February 2024.

    Comments: Paper accepted for SIGTYP (2024): Häuser, Luise; Jäger, Gerhard; List, Johann-Mattis; Rama, Taraka; and Stamatakis, Alexandros (2024): Are sounds sound for phylogenetic reconstruction? In: Proceedings of the 6th Workshop on Research in Computational Linguistic Typology and Multilingual NLP (SIGTYP 2024)

  3. arXiv:2204.05056  [pdf, ps, other

    cs.CL

    What do complexity measures measure? Correlating and validating corpus-based measures of morphological complexity

    Authors: Çağrı Çöltekin, Taraka Rama

    Abstract: We present an analysis of eight measures used for quantifying morphological complexity of natural languages. The measures we study are corpus-based measures of morphological complexity with varying requirements for corpus annotation. We present similarities and differences between these measures visually and through correlation analyses, as well as their relation to the relevant typological variab… ▽ More

    Submitted 11 April, 2022; originally announced April 2022.

    Comments: Submitted to Linguistics Vanguard

  4. arXiv:2102.12971  [pdf, other

    cs.CL

    Are pre-trained text representations useful for multilingual and multi-dimensional language proficiency modeling?

    Authors: Taraka Rama, Sowmya Vajjala

    Abstract: Development of language proficiency models for non-native learners has been an active area of interest in NLP research for the past few years. Although language proficiency is multidimensional in nature, existing research typically considers a single "overall proficiency" while building models. Further, existing approaches also considers only one language at a time. This paper describes our experi… ▽ More

    Submitted 25 February, 2021; originally announced February 2021.

    Comments: 10 pages

  5. arXiv:2011.02070  [pdf, other

    cs.CL

    Probing Multilingual BERT for Genetic and Typological Signals

    Authors: Taraka Rama, Lisa Beinborn, Steffen Eger

    Abstract: We probe the layers in multilingual BERT (mBERT) for phylogenetic and geographic language signals across 100 languages and compute language distances based on the mBERT representations. We 1) employ the language distances to infer and evaluate language trees, finding that they are close to the reference family tree in terms of quartet tree distance, 2) perform distance matrix regression analysis,… ▽ More

    Submitted 3 November, 2020; originally announced November 2020.

    Comments: COLING 2020

  6. arXiv:1809.04838  [pdf, ps, other

    cs.CL

    Tübingen-Oslo system: Linear regression works the best at Predicting Current and Future Psychological Health from Childhood Essays in the CLPsych 2018 Shared Task

    Authors: Çağrı Çöltekin, Taraka Rama

    Abstract: This paper describes our efforts in predicting current and future psychological health from childhood essays within the scope of the CLPsych-2018 Shared Task. We experimented with a number of different models, including recurrent and convolutional networks, Poisson regression, support vector regression, and L1 and L2 regularized linear regression. We obtained the best results on the training/devel… ▽ More

    Submitted 13 September, 2018; originally announced September 2018.

  7. arXiv:1805.03645  [pdf, other

    cs.CL

    Three tree priors and five datasets: A study of the effect of tree priors in Indo-European phylogenetics

    Authors: Taraka Rama

    Abstract: The age of the root of the Indo-European language family has received much attention since the application of Bayesian phylogenetic methods by Gray and Atkinson(2003). The root age of the Indo-European family has tended to decrease from an age that supported the Anatolian origin hypothesis to an age that supports the Steppe origin hypothesis with the application of new models (Chang et al., 2015).… ▽ More

    Submitted 9 May, 2018; originally announced May 2018.

  8. arXiv:1804.06636  [pdf, other

    cs.CL

    Experiments with Universal CEFR Classification

    Authors: Sowmya Vajjala, Taraka Rama

    Abstract: The Common European Framework of Reference (CEFR) guidelines describe language proficiency of learners on a scale of 6 levels. While the description of CEFR guidelines is generic across languages, the development of automated proficiency classification systems for different languages follow different approaches. In this paper, we explore universal CEFR classification using domain-specific and doma… ▽ More

    Submitted 18 April, 2018; originally announced April 2018.

    Comments: to appear in the proceedings of The 13th Workshop on Innovative Use of NLP for Building Educational Applications

  9. arXiv:1804.05416  [pdf, ps, other

    cs.CL

    Are Automatic Methods for Cognate Detection Good Enough for Phylogenetic Reconstruction in Historical Linguistics?

    Authors: Taraka Rama, Johann-Mattis List, Johannes Wahle, Gerhard Jäger

    Abstract: We evaluate the performance of state-of-the-art algorithms for automatic cognate detection by comparing how useful automatically inferred cognates are for the task of phylogenetic inference compared to classical manually annotated cognate sets. Our findings suggest that phylogenies inferred from automated cognate sets come close to phylogenies inferred from expert-annotated ones, although on avera… ▽ More

    Submitted 15 April, 2018; originally announced April 2018.

  10. arXiv:1702.04938  [pdf, other

    cs.CL

    Fast and unsupervised methods for multilingual cognate clustering

    Authors: Taraka Rama, Johannes Wahle, Pavel Sofroniev, Gerhard Jäger

    Abstract: In this paper we explore the use of unsupervised methods for detecting cognates in multilingual word lists. We use online EM to train sound segment similarity weights for computing similarity between two words. We tested our online systems on geographically spread sixteen different language groups of the world and show that the Online PMI system (Pointwise Mutual Information) outperforms a HMM bas… ▽ More

    Submitted 16 February, 2017; originally announced February 2017.

  11. arXiv:1610.06053  [pdf, ps, other

    cs.CL

    Chinese Restaurant Process for cognate clustering: A threshold free approach

    Authors: Taraka Rama

    Abstract: In this paper, we introduce a threshold free approach, motivated from Chinese Restaurant Process, for the purpose of cognate clustering. We show that our approach yields similar results to a linguistically motivated cognate clustering system known as LexStat. Our Chinese Restaurant Process system is fast and does not require any threshold and can be applied to any language family of the world.

    Submitted 19 October, 2016; originally announced October 2016.

  12. arXiv:1605.05172  [pdf, other

    cs.CL

    Siamese convolutional networks based on phonetic features for cognate identification

    Authors: Taraka Rama

    Abstract: In this paper, we explore the use of convolutional networks (ConvNets) for the purpose of cognate identification. We compare our architecture with binary classifiers based on string similarity measures on different language families. Our experiments show that convolutional networks achieve competitive results across concepts and across language families at the task of cognate identification.

    Submitted 2 July, 2016; v1 submitted 17 May, 2016; originally announced May 2016.

  13. arXiv:1409.0314  [pdf, ps, other

    cs.CL

    Empirical Evaluation of Tree distances for Parser Evaluation

    Authors: Taraka Rama

    Abstract: In this empirical study, I compare various tree distance measures -- originally developed in computational biology for the purpose of tree comparison -- for the purpose of parser evaluation. I will control for the parser setting by comparing the automatically generated parse trees from the state-of-the-art parser Charniak, 2000) with the gold-standard parse trees. The article describes two differe… ▽ More

    Submitted 2 September, 2014; v1 submitted 1 September, 2014; originally announced September 2014.

    Comments: Submitted to satisfy partial requirements for Statistical Parsing course

  14. arXiv:1408.2359  [pdf, other

    cs.CL

    Gap-weighted subsequences for automatic cognate identification and phylogenetic inference

    Authors: Taraka Rama

    Abstract: In this paper, we describe the problem of cognate identification and its relation to phylogenetic inference. We introduce subsequence based features for discriminating cognates from non-cognates. We show that subsequence based features perform better than the state-of-the-art string similarity measures for the purpose of cognate identification. We use the cognate judgments for the purpose of phylo… ▽ More

    Submitted 22 August, 2014; v1 submitted 11 August, 2014; originally announced August 2014.

  15. arXiv:1401.4869  [pdf

    cs.CL cs.AI

    Does Syntactic Knowledge help English-Hindi SMT?

    Authors: Taraka Rama, Karthik Gali, Avinesh PVS

    Abstract: In this paper we explore various parameter settings of the state-of-art Statistical Machine Translation system to improve the quality of the translation for a `distant' language pair like English-Hindi. We proposed new techniques for efficient reordering. A slight improvement over the baseline is reported using these techniques. We also show that a simple pre-processing step can improve the qualit… ▽ More

    Submitted 20 January, 2014; originally announced January 2014.

  16. arXiv:1401.0794  [pdf, other

    cs.CL stat.CO

    Properties of phoneme N -grams across the world's language families

    Authors: Taraka Rama, Lars Borin

    Abstract: In this article, we investigate the properties of phoneme N-grams across half of the world's languages. We investigate if the sizes of three different N-gram distributions of the world's language families obey a power law. Further, the N-gram distributions of language families parallel the sizes of the families, which seem to obey a power law distribution. The correlation between N-gram distributi… ▽ More

    Submitted 4 January, 2014; originally announced January 2014.

  17. arXiv:1401.0708  [pdf

    cs.CL cs.AI

    Quantitative methods for Phylogenetic Inference in Historical Linguistics: An experimental case study of South Central Dravidian

    Authors: Taraka Rama, Sudheer Kolachina, Lakshmi Bai B

    Abstract: In this paper we examine the usefulness of two classes of algorithms Distance Methods, Discrete Character Methods (Felsenstein and Felsenstein 2003) widely used in genetics, for predicting the family relationships among a set of related languages and therefore, diachronic language change. Applying these algorithms to the data on the numbers of shared cognates- with-change and changed as well as un… ▽ More

    Submitted 3 January, 2014; originally announced January 2014.

    Journal ref: Indian Linguistics, Volume 70, 2009