-
Critical mobility in policy making for epidemic containment
Authors:
Jesús A. Moreno López,
Sandro Meloni,
Jose J. Ramasco
Abstract:
When considering airborne epidemic spreading in social systems, a natural connection arises between mobility and epidemic contacts. As individuals travel, possibilities to encounter new people either at the final destination or during the transportation process appear. Such contacts can lead to new contagion events. In fact, mobility has been a crucial target for early non-pharmaceutical containme…
▽ More
When considering airborne epidemic spreading in social systems, a natural connection arises between mobility and epidemic contacts. As individuals travel, possibilities to encounter new people either at the final destination or during the transportation process appear. Such contacts can lead to new contagion events. In fact, mobility has been a crucial target for early non-pharmaceutical containment measures against the recent COVID-19 pandemic, with a degree of intensity ranging from public transportation line closures to regional, city or even home confinements. Nonetheless, quantitative knowledge on the relationship between mobility-contagions and, consequently, on the efficiency of containment measures remains elusive. Here we introduce an agent-based model with a simple interaction between mobility and contacts. Despite its simplicity our model shows the emergence of a critical mobility level, inducing major outbreaks when surpassed. We explore the interplay between mobility restrictions and the infection in recent intervention policies seen across many countries, and how interventions in the form of closures triggered by incidence rates can guide the epidemic into an oscillatory regime with recurrent waves. We consider how the different interventions impact societal well-being, the economy and the population. Finally, we propose a mitigation framework based on the critical nature of mobility in an epidemic, able to suppress incidence and oscillations at will, preventing extreme incidence peaks with potential to saturate health care resources.
△ Less
Submitted 9 May, 2024; v1 submitted 8 February, 2024;
originally announced February 2024.
-
A generalized vector-field framework for mobility
Authors:
Erjian Liu,
Mattia Mazzoli,
Xiao-Yong Yan,
Jose J. Ramasco
Abstract:
Trip flow between areas is a fundamental metric for human mobility research. Given its identification with travel demand and its relevance for transportation and urban planning, many models have been developed for its estimation. These models focus on flow intensity, disregarding the information provided by the local mobility orientation. A field-theoretic approach can overcome this issue and hand…
▽ More
Trip flow between areas is a fundamental metric for human mobility research. Given its identification with travel demand and its relevance for transportation and urban planning, many models have been developed for its estimation. These models focus on flow intensity, disregarding the information provided by the local mobility orientation. A field-theoretic approach can overcome this issue and handling both intensity and direction at once. Here we propose a general vector-field representation starting from individuals' trajectories valid for any type of mobility. By introducing four models of spatial exploration, we show how individuals' elections determine the mesoscopic properties of the mobility field. Distance optimization in long displacements and random-like local exploration are necessary to reproduce empirical field features observed in Chinese logistic data and in New York City Foursquare check-ins. Our framework is an essential tool to capture hidden symmetries in mesoscopic urban mobility, it establishes a benchmark to test the validity of mobility models and opens the doors to the use of field theory in a wide spectrum of applications.
△ Less
Submitted 4 September, 2023;
originally announced September 2023.
-
When Dialects Collide: How Socioeconomic Mixing Affects Language Use
Authors:
Thomas Louf,
José J. Ramasco,
David Sánchez,
Márton Karsai
Abstract:
The socioeconomic background of people and how they use standard forms of language are not independent, as demonstrated in various sociolinguistic studies. However, the extent to which these correlations may be influenced by the mixing of people from different socioeconomic classes remains relatively unexplored from a quantitative perspective. In this work we leverage geotagged tweets and transfer…
▽ More
The socioeconomic background of people and how they use standard forms of language are not independent, as demonstrated in various sociolinguistic studies. However, the extent to which these correlations may be influenced by the mixing of people from different socioeconomic classes remains relatively unexplored from a quantitative perspective. In this work we leverage geotagged tweets and transferable computational methods to map deviations from standard English on a large scale, in seven thousand administrative areas of England and Wales. We combine these data with high-resolution income maps to assign a proxy socioeconomic indicator to home-located users. Strikingly, across eight metropolitan areas we find a consistent pattern suggesting that the more different socioeconomic classes mix, the less interdependent the frequency of their departures from standard grammar and their income become. Further, we propose an agent-based model of linguistic variety adoption that sheds light on the mechanisms that produce the observations seen in the data.
△ Less
Submitted 10 July, 2025; v1 submitted 19 July, 2023;
originally announced July 2023.
-
American cultural regions mapped through the lexical analysis of social media
Authors:
Thomas Louf,
Bruno Gonçalves,
Jose J. Ramasco,
David Sanchez,
Jack Grieve
Abstract:
Cultural areas represent a useful concept that cross-fertilizes diverse fields in social sciences. Knowledge of how humans organize and relate their ideas and behavior within a society helps to understand their actions and attitudes towards different issues. However, the selection of common traits that shape a cultural area is somewhat arbitrary. What is needed is a method that can leverage the ma…
▽ More
Cultural areas represent a useful concept that cross-fertilizes diverse fields in social sciences. Knowledge of how humans organize and relate their ideas and behavior within a society helps to understand their actions and attitudes towards different issues. However, the selection of common traits that shape a cultural area is somewhat arbitrary. What is needed is a method that can leverage the massive amounts of data coming online, especially through social media, to identify cultural regions without ad-hoc assumptions, biases or prejudices. This work takes a crucial step in this direction by introducing a method to infer cultural regions based on the automatic analysis of large datasets from microblogging posts. The approach presented here is based on the principle that cultural affiliation can be inferred from the topics that people discuss among themselves. Specifically, regional variations in written discourse are measured in American social media. From the frequency distributions of content words in geotagged Tweets, the regional hotspots of words' usage are found, and from there, principal components of regional variation are derived. Through a hierarchical clustering of the data in this lower-dimensional space, this method yields clear cultural areas and the topics of discussion that define them. It uncovers a manifest North-South separation, which is primarily influenced by the African American culture, and further contiguous (East-West) and non-contiguous divisions that provide a comprehensive picture of today's cultural areas in the US.
△ Less
Submitted 18 April, 2023; v1 submitted 16 August, 2022;
originally announced August 2022.
-
Capturing the diversity of multilingual societies
Authors:
Thomas Louf,
David Sanchez,
Jose J. Ramasco
Abstract:
Cultural diversity encoded within languages of the world is at risk, as many languages have become endangered in the last decades in a context of growing globalization. To preserve this diversity, it is first necessary to understand what drives language extinction, and which mechanisms might enable coexistence. Here, we study language shift mechanisms using theoretical and data-driven perspectives…
▽ More
Cultural diversity encoded within languages of the world is at risk, as many languages have become endangered in the last decades in a context of growing globalization. To preserve this diversity, it is first necessary to understand what drives language extinction, and which mechanisms might enable coexistence. Here, we study language shift mechanisms using theoretical and data-driven perspectives. A large-scale empirical analysis of multilingual societies using Twitter and census data yields a wide diversity of spatial patterns of language coexistence. It ranges from a mixing of language speakers to segregation with multilinguals on the boundaries of disjoint linguistic domains. To understand how these different states can emerge and, especially, become stable, we propose a model in which language coexistence is reached when learning the other language is facilitated and when bilinguals favor the use of the endangered language. Simulations carried out in a metapopulation framework highlight the importance of spatial interactions arising from people mobility to explain the stability of a mixed state or the presence of a boundary between two linguistic regions. Further, we find that the history of languages is critical to understand their present state.
△ Less
Submitted 7 October, 2022; v1 submitted 6 May, 2021;
originally announced May 2021.
-
The world-wide waste web
Authors:
Johann H. Martínez,
Sergi Romero,
José J. Ramasco,
Ernesto Estrada
Abstract:
Countries globally trade with tons of waste materials every year, some of which are highly hazardous. This trade admits a network representation of the world-wide waste web, with countries as vertices and flows as directed weighted edges. Here we investigate the main properties of this network by tracking 108 categories of wastes interchanged in the period 2001-2019. Although, most of the hazardou…
▽ More
Countries globally trade with tons of waste materials every year, some of which are highly hazardous. This trade admits a network representation of the world-wide waste web, with countries as vertices and flows as directed weighted edges. Here we investigate the main properties of this network by tracking 108 categories of wastes interchanged in the period 2001-2019. Although, most of the hazardous waste was traded between developed nations, a disproportionate asymmetry existed in the flow from developed to developing countries. Using a dynamical model, we simulate how waste stress propagates through the network and affects the countries. We identify 28 countries with low Environmental Performance Index that are at high risk of waste congestion. Therefore, they are at threat of improper handling and disposal of hazardous waste. We find evidence of pollution by heavy metals, by volatile organic compounds and/or by persistent organic pollutants, which are used as chemical fingerprints, due to the improper handling of waste in several of these countries.
△ Less
Submitted 14 March, 2022; v1 submitted 12 April, 2021;
originally announced April 2021.
-
Field theory for recurrent mobility
Authors:
Mattia Mazzoli,
Alex Molas,
Aleix Bassolas,
Maxime Lenormand,
Pere Colet,
Jose J. Ramasco
Abstract:
Understanding human mobility is crucial for applications such as forecasting epidemic spreading, planning transport infrastructure and urbanism in general. While, traditionally, mobility information has been collected via surveys, the pervasive adoption of mobile technologies has brought a wealth of (real time) data. The easy access to this information opens the door to study theoretical questions…
▽ More
Understanding human mobility is crucial for applications such as forecasting epidemic spreading, planning transport infrastructure and urbanism in general. While, traditionally, mobility information has been collected via surveys, the pervasive adoption of mobile technologies has brought a wealth of (real time) data. The easy access to this information opens the door to study theoretical questions so far unexplored. In this work, we show for a series of worldwide cities that commuting daily flows can be mapped into a well behaved vector field, fulfilling the divergence theorem and which is, besides, irrotational. This property allows us to define a potential for the field that can become a major instrument to determine separate mobility basins and discern contiguous urban areas. We also show that empirical fluxes and potentials can be well reproduced and analytically characterized using the so-called gravity model, while other models based on intervening opportunities have serious difficulties.
△ Less
Submitted 30 August, 2019;
originally announced August 2019.
-
Migrant mobility flows characterized with digital data
Authors:
Mattia Mazzoli,
Boris Diechtiareff,
Antonia Tugores,
Willian Wives,
Natalia Adler,
Pere Colet,
Jose J. Ramasco
Abstract:
Monitoring migration flows is crucial to respond to humanitarian crisis and to design efficient policies. This information usually comes from surveys and border controls, but timely accessibility and methodological concerns reduce its usefulness. Here, we propose a method to detect migration flows worldwide using geolocated Twitter data. We focus on the migration crisis in Venezuela and show that…
▽ More
Monitoring migration flows is crucial to respond to humanitarian crisis and to design efficient policies. This information usually comes from surveys and border controls, but timely accessibility and methodological concerns reduce its usefulness. Here, we propose a method to detect migration flows worldwide using geolocated Twitter data. We focus on the migration crisis in Venezuela and show that the calculated flows are consistent with official statistics at country level. Our method is versatile and far-reaching, as it can be used to study different features of migration as preferred routes, settlement areas, mobility through several countries, spatial integration in cities, etc. It provides finer geographical and temporal resolutions, allowing the exploration of issues not contemplated in official records. It is our hope that these new sources of information can complement official ones, helping authorities and humanitarian organizations to better assess when and where to intervene on the ground.
△ Less
Submitted 7 August, 2019;
originally announced August 2019.
-
Scaling in the recovery of urban transportation systems from special events
Authors:
Aleix Bassolas,
Riccardo Gallotti,
Fabio Lamanna,
Maxime Lenormand,
Jose J. Ramasco
Abstract:
Public transportation is a fundamental infrastructure for the daily mobility in cities. Although its capacity is prepared for the usual demand, congestion may rise when huge crowds concentrate in special events such as massive demonstrations, concerts or sport events. In this work, we study the resilience and recovery of public transportation networks from massive gatherings by means of a stylized…
▽ More
Public transportation is a fundamental infrastructure for the daily mobility in cities. Although its capacity is prepared for the usual demand, congestion may rise when huge crowds concentrate in special events such as massive demonstrations, concerts or sport events. In this work, we study the resilience and recovery of public transportation networks from massive gatherings by means of a stylized model mimicking the mobility of individuals through the multilayer transportation network. We focus on the delays produced by the congestion in the trips of both event participants and of other citizens doing their usual traveling in the background. Our model can be solved analytically for regular lattices showing that the average delay scales with the number of event participants with an exponent equal to the inverse of the lattice dimension. We then switch to real transportation networks of eight worldwide cities, and observe that there is a whole range of exponents depending on where the event is located. These exponents are distributed around 1/2, which indicates that most of the local structure of the network is two dimensional. Yet, some of the exponents are below (above) that value, implying a local dimension higher (lower) than 2 as a consequence of the multimodality and multifractality of transportation networks. In fact, these exponents can be also obtained from the scaling of the capacity with the distance from the event. Overall, our methodology allows to dynamically probe the local dimensionality of a transportation network and identify the most vulnerable spots in cities for the celebration of massive events.
△ Less
Submitted 19 June, 2019;
originally announced June 2019.
-
Mobile phone records to feed activity-based travel demand models: MATSim for studying a cordon toll policy in Barcelona
Authors:
Aleix Bassolas,
Jose J. Ramasco,
Ricardo Herranz,
Oliva G. Cantu-Ros
Abstract:
Activity-based models appeared as an answer to the limitations of the traditional trip-based and tour-based four-stage models. The fundamental assumption of activity-based models is that travel demand is originated from people performing their daily activities. This is why they include a consistent representation of time, of the persons and households, time-dependent routing, and microsimulation o…
▽ More
Activity-based models appeared as an answer to the limitations of the traditional trip-based and tour-based four-stage models. The fundamental assumption of activity-based models is that travel demand is originated from people performing their daily activities. This is why they include a consistent representation of time, of the persons and households, time-dependent routing, and microsimulation of travel demand and traffic. In spite of their potential to simulate traffic demand management policies, their practical application is still limited. One of the main reasons is that these models require a huge amount of very detailed input data hard to get with surveys. However, the pervasive use of mobile devices has brought a valuable new source of data. The work presented here has a twofold objective: first, to demonstrate the capability of mobile phone records to feed activity-based transport models, and, second, to assert the advantages of using activity-based models to estimate the effects of traffic demand management policies. Activity diaries for the metropolitan area of Barcelona are reconstructed from mobile phone records. This information is then employed as input for building a transport MATSim model of the city. The model calibration and validation process proves the quality of the activity diaries obtained. The possible impacts of a cordon toll policy applied to two different areas of the city and at different times of the day is then studied. Our results show the way in which the modal share is modified in each of the considered scenario. The possibility of evaluating the effects of the policy at both aggregated and traveller level, together with the ability of the model to capture policy impacts beyond the cordon toll area confirm the advantages of activity-based models for the evaluation of traffic demand management policies.
△ Less
Submitted 16 March, 2018;
originally announced March 2018.
-
Mapping the Americanization of English in Space and Time
Authors:
Bruno Gonçalves,
Lucía Loureiro-Porto,
José J. Ramasco,
David Sánchez
Abstract:
As global political preeminence gradually shifted from the United Kingdom to the United States, so did the capacity to culturally influence the rest of the world. In this work, we analyze how the world-wide varieties of written English are evolving. We study both the spatial and temporal variations of vocabulary and spelling of English using a large corpus of geolocated tweets and the Google Books…
▽ More
As global political preeminence gradually shifted from the United Kingdom to the United States, so did the capacity to culturally influence the rest of the world. In this work, we analyze how the world-wide varieties of written English are evolving. We study both the spatial and temporal variations of vocabulary and spelling of English using a large corpus of geolocated tweets and the Google Books datasets corresponding to books published in the US and the UK. The advantage of our approach is that we can address both standard written language (Google Books) and the more colloquial forms of microblogging messages (Twitter). We find that American English is the dominant form of English outside the UK and that its influence is felt even within the UK borders. Finally, we analyze how this trend has evolved over time and the impact that some cultural events have had in shaping it.
△ Less
Submitted 28 May, 2018; v1 submitted 3 July, 2017;
originally announced July 2017.
-
Immigrant community integration in world cities
Authors:
Fabio Lamanna,
Maxime Lenormand,
María Henar Salas-Olmedo,
Gustavo Romanillos,
Bruno Gonçalves,
José J. Ramasco
Abstract:
As a consequence of the accelerated globalization process, today major cities all over the world are characterized by an increasing multiculturalism. The integration of immigrant communities may be affected by social polarization and spatial segregation. How are these dynamics evolving over time? To what extent the different policies launched to tackle these problems are working? These are critica…
▽ More
As a consequence of the accelerated globalization process, today major cities all over the world are characterized by an increasing multiculturalism. The integration of immigrant communities may be affected by social polarization and spatial segregation. How are these dynamics evolving over time? To what extent the different policies launched to tackle these problems are working? These are critical questions traditionally addressed by studies based on surveys and census data. Such sources are safe to avoid spurious biases, but the data collection becomes an intensive and rather expensive work. Here, we conduct a comprehensive study on immigrant integration in 53 world cities by introducing an innovative approach: an analysis of the spatio-temporal communication patterns of immigrant and local communities based on language detection in Twitter and on novel metrics of spatial integration. We quantify the "Power of Integration" of cities --their capacity to spatially integrate diverse cultures-- and characterize the relations between different cultures when acting as hosts or immigrants.
△ Less
Submitted 14 March, 2018; v1 submitted 3 November, 2016;
originally announced November 2016.
-
Is spatial information in ICT data reliable?
Authors:
Maxime Lenormand,
Thomas Louail,
Marc Barthelemy,
José J. Ramasco
Abstract:
An increasing number of human activities are studied using data produced by individuals' ICT devices. In particular, when ICT data contain spatial information, they represent an invaluable source for analyzing urban dynamics. However, there have been relatively few contributions investigating the robustness of this type of results against fluctuations of data characteristics. Here, we present a st…
▽ More
An increasing number of human activities are studied using data produced by individuals' ICT devices. In particular, when ICT data contain spatial information, they represent an invaluable source for analyzing urban dynamics. However, there have been relatively few contributions investigating the robustness of this type of results against fluctuations of data characteristics. Here, we present a stability analysis of higher-level information extracted from mobile phone data passively produced during an entire year by 9 million individuals in Senegal. We focus on two information-retrieval tasks: (a) the identification of land use in the region of Dakar from the temporal rhythms of the communication activity; (b) the identification of home and work locations of anonymized individuals, which enable to construct Origin-Destination (OD) matrices of commuting flows. Our analysis reveal that the uncertainty of results highly depends on the sample size, the scale and the period of the year at which the data were gathered. Nevertheless, the spatial distributions of land use computed for different samples are remarkably robust: on average, we observe more than 75% of shared surface area between the different spatial partitions when considering activity of at least 100,000 users whatever the scale. The OD matrix is less stable and depends on the scale with a share of at least 75% of commuters in common when considering all types of flows constructed from the home-work locations of 100,000 users. For both tasks, better results can be obtained at larger levels of aggregation or by considering more users. These results confirm that ICT data are very useful sources for the spatial analysis of urban systems, but that their reliability should in general be tested more thoroughly.
△ Less
Submitted 23 November, 2017; v1 submitted 12 September, 2016;
originally announced September 2016.
-
Crowdsourcing the Robin Hood effect in cities
Authors:
Thomas Louail,
Maxime Lenormand,
Juan Murillo Arias,
José J. Ramasco
Abstract:
Socioeconomic inequalities in cities are embedded in space and result in neighborhood effects, whose harmful consequences have proved very hard to counterbalance efficiently by planning policies alone. Considering redistribution of money flows as a first step toward improved spatial equity, we study a bottom-up approach that would rely on a slight evolution of shopping mobility practices. Building…
▽ More
Socioeconomic inequalities in cities are embedded in space and result in neighborhood effects, whose harmful consequences have proved very hard to counterbalance efficiently by planning policies alone. Considering redistribution of money flows as a first step toward improved spatial equity, we study a bottom-up approach that would rely on a slight evolution of shopping mobility practices. Building on a database of anonymized credit card transactions in Madrid and Barcelona, we quantify the mobility effort required to reach a reference situation where commercial income is evenly shared among neighborhoods. The redirections of shopping trips preserve key properties of human mobility, including travel distances. Surprisingly, for both cities only a small fraction ($\sim 5 \%$) of trips need to be altered to reach equity situations, improving even other sustainability indicators. The method could be implemented in mobile applications that would assist individuals in reshaping their shopping practices, to promote the spatial redistribution of opportunities in the city.
△ Less
Submitted 9 June, 2017; v1 submitted 28 April, 2016;
originally announced April 2016.
-
Dynamics on networks: competition of temporal and topological correlations
Authors:
Oriol Artime,
Jose J. Ramasco,
Maxi San Miguel
Abstract:
Links in many real-world networks activate and deactivate in correspondence to the sporadic interactions between the elements of the system. The activation patterns may be irregular or bursty and play an important role on the dynamics of processes taking place in the network. Information or disease spreading in networks are paradigmatic examples of this situation. Besides burstiness, several corre…
▽ More
Links in many real-world networks activate and deactivate in correspondence to the sporadic interactions between the elements of the system. The activation patterns may be irregular or bursty and play an important role on the dynamics of processes taking place in the network. Information or disease spreading in networks are paradigmatic examples of this situation. Besides burstiness, several correlations may appear in the process of link activation: memory effects imply temporal correlations, but also the existence of communities in the network may mediate the activation patterns of internal an external links. Here we study the competition of topological and temporal correlations in link activation and how they affect the dynamics of systems running on the network. Interestingly, both types of correlations by separate have opposite effects: one (topological) delays the dynamics of processes on the network, while the other (temporal) accelerates it. When they occur together, our results show that the direction and intensity of the final outcome depends on the competition in a non trivial way.
△ Less
Submitted 1 February, 2017; v1 submitted 14 April, 2016;
originally announced April 2016.
-
Dynamical leaps due to microscopic changes in multiplex networks
Authors:
Marina Diakonova,
Jose J. Ramasco,
Victor M. Eguiluz
Abstract:
Recent developments of the multiplex paradigm included efforts to understand the role played by the presence of several layers on the dynamics of processes running on these networks. The possible existence of new phenomena associated to the richer topology has been discussed and examples of these differences have been systematically searched. Here, we show that the interconnectivity of the layers…
▽ More
Recent developments of the multiplex paradigm included efforts to understand the role played by the presence of several layers on the dynamics of processes running on these networks. The possible existence of new phenomena associated to the richer topology has been discussed and examples of these differences have been systematically searched. Here, we show that the interconnectivity of the layers may have an important impact on the speed of the dynamics run in the network and that microscopic changes such as the addition of one single inter-layer link can notably affect the arrival at a global stationary state. As a practical verification, these results obtained with spectral techniques are confirmed with a Kuramoto dynamics for which the synchronization consistently delays after the addition of single inter-layer links.
△ Less
Submitted 4 April, 2016;
originally announced April 2016.
-
Touristic site attractiveness seen through Twitter
Authors:
Aleix Bassolas,
Maxime Lenormand,
Antònia Tugores,
Bruno Gonçalves,
José J. Ramasco
Abstract:
Tourism is becoming a significant contributor to medium and long range travels in an increasingly globalized world. Leisure traveling has an important impact on the local and global economy as well as on the environment. The study of touristic trips is thus raising a considerable interest. In this work, we apply a method to assess the attractiveness of 20 of the most popular touristic sites worldw…
▽ More
Tourism is becoming a significant contributor to medium and long range travels in an increasingly globalized world. Leisure traveling has an important impact on the local and global economy as well as on the environment. The study of touristic trips is thus raising a considerable interest. In this work, we apply a method to assess the attractiveness of 20 of the most popular touristic sites worldwide using geolocated tweets as a proxy for human mobility. We first rank the touristic sites based on the spatial distribution of the visitors' place of residence. The Taj Mahal, the Pisa Tower and the Eiffel Tower appear consistently in the top 5 in these rankings. We then pass to a coarser scale and classify the travelers by country of residence. Touristic site's visiting figures are then studied by country of residence showing that the Eiffel Tower, Times Square and the London Tower welcome the majority of the visitors of each country. Finally, we build a network linking sites whenever a user has been detected in more than one site. This allow us to unveil relations between touristic sites and find which ones are more tightly interconnected.
△ Less
Submitted 25 March, 2016; v1 submitted 28 January, 2016;
originally announced January 2016.
-
Persistence in voting behavior: stronghold dynamics in elections
Authors:
Toni Pérez,
Juan Fernández-Gracia,
Jose J. Ramasco,
Víctor M. Eguíluz
Abstract:
Influence among individuals is at the core of collective social phenomena such as the dissemination of ideas, beliefs or behaviors, social learning and the diffusion of innovations. Different mechanisms have been proposed to implement inter-agent influence in social models from the voter model, to majority rules, to the Granoveter model. Here we advance in this direction by confronting the recentl…
▽ More
Influence among individuals is at the core of collective social phenomena such as the dissemination of ideas, beliefs or behaviors, social learning and the diffusion of innovations. Different mechanisms have been proposed to implement inter-agent influence in social models from the voter model, to majority rules, to the Granoveter model. Here we advance in this direction by confronting the recently introduced Social Influence and Recurrent Mobility (SIRM) model, that reproduces generic features of vote-shares at different geographical levels, with data in the US presidential elections. Our approach incorporates spatial and population diversity as inputs for the opinion dynamics while individuals' mobility provides a proxy for social context, and peer imitation accounts for social influence. The model captures the observed stationary background fluctuations in the vote-shares across counties. We study the so-called political strongholds, i.e., locations where the votes-shares for a party are systematically higher than average. A quantitative definition of a stronghold by means of persistence in time of fluctuations in the voting spatial distribution is introduced, and results from the US Presidential Elections during the period 1980-2012 are analyzed within this framework. We compare electoral results with simulations obtained with the SIRM model finding a good agreement both in terms of the number and the location of strongholds. The strongholds duration is also systematically characterized in the SIRM model. The results compare well with the electoral results data revealing an exponential decay in the persistence of the strongholds with time.
△ Less
Submitted 23 March, 2015;
originally announced March 2015.
-
Human diffusion and city influence
Authors:
Maxime Lenormand,
Bruno Gonçalves,
Antònia Tugores,
José J. Ramasco
Abstract:
Cities are characterized by concentrating population, economic activity and services. However, not all cities are equal and a natural hierarchy at local, regional or global scales spontaneously emerges. In this work, we introduce a method to quantify city influence using geolocated tweets to characterize human mobility. Rome and Paris appear consistently as the cities attracting most diverse visit…
▽ More
Cities are characterized by concentrating population, economic activity and services. However, not all cities are equal and a natural hierarchy at local, regional or global scales spontaneously emerges. In this work, we introduce a method to quantify city influence using geolocated tweets to characterize human mobility. Rome and Paris appear consistently as the cities attracting most diverse visitors. The ratio between locals and non-local visitors turns out to be fundamental for a city to truly be global. Focusing only on urban residents' mobility flows, a city to city network can be constructed. This network allows us to analyze centrality measures at different scales. New York and London play a predominant role at the global scale, while urban rankings suffer substantial changes if the focus is set at a regional level.
△ Less
Submitted 15 July, 2015; v1 submitted 30 January, 2015;
originally announced January 2015.
-
Uncovering the spatial structure of mobility networks
Authors:
Thomas Louail,
Maxime Lenormand,
Miguel Picornell,
Oliva García Cantú,
Ricardo Herranz,
Enrique Frias-Martinez,
José J. Ramasco,
Marc Barthelemy
Abstract:
The extraction of a clear and simple footprint of the structure of large, weighted and directed networks is a general problem that has many applications. An important example is given by origin-destination matrices which contain the complete information on commuting flows, but are difficult to analyze and compare. We propose here a versatile method which extracts a coarse-grained signature of mobi…
▽ More
The extraction of a clear and simple footprint of the structure of large, weighted and directed networks is a general problem that has many applications. An important example is given by origin-destination matrices which contain the complete information on commuting flows, but are difficult to analyze and compare. We propose here a versatile method which extracts a coarse-grained signature of mobility networks, under the form of a $2\times 2$ matrix that separates the flows into four categories. We apply this method to origin-destination matrices extracted from mobile phone data recorded in thirty-one Spanish cities. We show that these cities essentially differ by their proportion of two types of flows: integrated (between residential and employment hotspots) and random flows, whose importance increases with city size. Finally the method allows to determine categories of networks, and in the mobility case to classify cities according to their commuting structure.
△ Less
Submitted 21 January, 2015;
originally announced January 2015.
-
Influence of sociodemographic characteristics on human mobility
Authors:
Maxime Lenormand,
Thomas Louail,
Oliva G. Cantu-Ros,
Miguel Picornell,
Ricardo Herranz,
Juan Murillo Arias,
Marc Barthelemy,
Maxi San Miguel,
Jose J. Ramasco
Abstract:
Human mobility has been traditionally studied using surveys that deliver snapshots of population displacement patterns. The growing accessibility to ICT information from portable digital media has recently opened the possibility of exploring human behavior at high spatio-temporal resolutions. Mobile phone records, geolocated tweets, check-ins from Foursquare or geotagged photos, have contributed t…
▽ More
Human mobility has been traditionally studied using surveys that deliver snapshots of population displacement patterns. The growing accessibility to ICT information from portable digital media has recently opened the possibility of exploring human behavior at high spatio-temporal resolutions. Mobile phone records, geolocated tweets, check-ins from Foursquare or geotagged photos, have contributed to this purpose at different scales, from cities to countries, in different world areas. Many previous works lacked, however, details on the individuals' attributes such as age or gender. In this work, we analyze credit-card records from Barcelona and Madrid and by examining the geolocated credit-card transactions of individuals living in the two provinces, we find that the mobility patterns vary according to gender, age and occupation. Differences in distance traveled and travel purpose are observed between younger and older people, but, curiously, either between males and females of similar age. While mobility displays some generic features, here we show that sociodemographic characteristics play a relevant role and must be taken into account for mobility and epidemiological modelization.
△ Less
Submitted 25 May, 2015; v1 submitted 28 November, 2014;
originally announced November 2014.
-
Tweets on the road
Authors:
Maxime Lenormand,
Antònia Tugores,
Pere Colet,
José J. Ramasco
Abstract:
The pervasiveness of mobile devices, which is increasing daily, is generating a vast amount of geo-located data allowing us to gain further insights into human behaviors. In particular, this new technology enables users to communicate through mobile social media applications, such as Twitter, anytime and anywhere. Thus, geo-located tweets offer the possibility to carry out in-depth studies on huma…
▽ More
The pervasiveness of mobile devices, which is increasing daily, is generating a vast amount of geo-located data allowing us to gain further insights into human behaviors. In particular, this new technology enables users to communicate through mobile social media applications, such as Twitter, anytime and anywhere. Thus, geo-located tweets offer the possibility to carry out in-depth studies on human mobility. In this paper, we study the use of Twitter in transportation by identifying tweets posted from roads and rails in Europe between September 2012 and November 2013. We compute the percentage of highway and railway segments covered by tweets in 39 countries. The coverages are very different from country to country and their variability can be partially explained by differences in Twitter penetration rates. Still, some of these differences might be related to cultural factors regarding mobility habits and interacting socially online. Analyzing particular road sectors, our results show a positive correlation between the number of tweets on the road and the Average Annual Daily Traffic on highways in France and in the UK. Transport modality can be studied with these data as well, for which we discover very heterogeneous usage patterns across the continent.
△ Less
Submitted 20 May, 2014;
originally announced May 2014.
-
Cross-checking different sources of mobility information
Authors:
Maxime Lenormand,
Miguel Picornell,
Oliva G. Cantu-Ros,
Antonia Tugores,
Thomas Louail,
Ricardo Herranz,
Marc Barthelemy,
Enrique Frias-Martinez,
Jose J. Ramasco
Abstract:
The pervasive use of new mobile devices has allowed a better characterization in space and time of human concentrations and mobility in general. Besides its theoretical interest, describing mobility is of great importance for a number of practical applications ranging from the forecast of disease spreading to the design of new spaces in urban environments. While classical data sources, such as sur…
▽ More
The pervasive use of new mobile devices has allowed a better characterization in space and time of human concentrations and mobility in general. Besides its theoretical interest, describing mobility is of great importance for a number of practical applications ranging from the forecast of disease spreading to the design of new spaces in urban environments. While classical data sources, such as surveys or census, have a limited level of geographical resolution (e.g., districts, municipalities, counties are typically used) or are restricted to generic workdays or weekends, the data coming from mobile devices can be precisely located both in time and space. Most previous works have used a single data source to study human mobility patterns. Here we perform instead a cross-check analysis by comparing results obtained with data collected from three different sources: Twitter, census and cell phones. The analysis is focused on the urban areas of Barcelona and Madrid, for which data of the three types is available. We assess the correlation between the datasets on different aspects: the spatial distribution of people concentration, the temporal evolution of people density and the mobility patterns of individuals. Our results show that the three data sources are providing comparable information. Even though the representativeness of Twitter geolocated data is lower than that of mobile phone and census data, the correlations between the population density profiles and mobility patterns detected by the three datasets are close to one in a grid with cells of 2x2 and 1x1 square kilometers. This level of correlation supports the feasibility of interchanging the three data sources at the spatio-temporal scales considered.
△ Less
Submitted 26 August, 2014; v1 submitted 1 April, 2014;
originally announced April 2014.
-
Is the Voter Model a model for voters?
Authors:
Juan Fernández-Gracia,
Krzysztof Suchecki,
José J. Ramasco,
Maxi San Miguel,
Víctor M. Eguíluz
Abstract:
The voter model has been studied extensively as a paradigmatic opinion dynamics' model. However, its ability for modeling real opinion dynamics has not been addressed. We introduce a noisy voter model (accounting for social influence) with agents' recurrent mobility (as a proxy for social context), where the spatial and population diversity are taken as inputs to the model. We show that the dynami…
▽ More
The voter model has been studied extensively as a paradigmatic opinion dynamics' model. However, its ability for modeling real opinion dynamics has not been addressed. We introduce a noisy voter model (accounting for social influence) with agents' recurrent mobility (as a proxy for social context), where the spatial and population diversity are taken as inputs to the model. We show that the dynamics can be described as a noisy diffusive process that contains the proper anysotropic coupling topology given by population and mobility heterogeneity. The model captures statistical features of the US presidential elections as the stationary vote-share fluctuations across counties, and the long-range spatial correlations that decay logarithmically with the distance. Furthermore, it recovers the behavior of these properties when a real-space renormalization is performed by coarse-graining the geographical scale from county level through congressional districts and up to states. Finally, we analyze the role of the mobility range and the randomness in decision making which are consistent with the empirical observations.
△ Less
Submitted 16 June, 2014; v1 submitted 4 September, 2013;
originally announced September 2013.
-
Entangling mobility and interactions in social media
Authors:
Przemyslaw A. Grabowicz,
Jose J. Ramasco,
Bruno Goncalves,
Victor M. Eguiluz
Abstract:
Daily interactions naturally define social circles. Individuals tend to be friends with the people they spend time with and they choose to spend time with their friends, inextricably entangling physical location and social relationships. As a result, it is possible to predict not only someone's location from their friends' locations but also friendship from spatial and temporal co-occurrence. Whil…
▽ More
Daily interactions naturally define social circles. Individuals tend to be friends with the people they spend time with and they choose to spend time with their friends, inextricably entangling physical location and social relationships. As a result, it is possible to predict not only someone's location from their friends' locations but also friendship from spatial and temporal co-occurrence. While several models have been developed to separately describe mobility and the evolution of social networks, there is a lack of studies coupling social interactions and mobility. In this work, we introduce a new model that bridges this gap by explicitly considering the feedback of mobility on the formation of social ties. Data coming from three online social networks (Twitter, Gowalla and Brightkite) is used for validation. Our model reproduces various topological and physical properties of these networks such as: i) the size of the connected components, ii) the distance distribution between connected users, iii) the dependence of the reciprocity on the distance, iv) the variation of the social overlap and the clustering with the distance. Besides numerical simulations, a mean-field approach is also used to study analytically the main statistical features of the networks generated by the model. The robustness of the results to changes in the model parameters is explored, finding that a balance between friend visits and long-range random connections is essential to reproduce the geographical features of the empirical networks.
△ Less
Submitted 10 March, 2014; v1 submitted 19 July, 2013;
originally announced July 2013.
-
Characterization of delay propagation in the US air transportation network
Authors:
Pablo Fleurquin,
José J. Ramasco,
Victor M. Eguíluz
Abstract:
Complex networks provide a suitable framework to characterize air traffic. Previous works described the world air transport network as a graph where direct flights are edges and commercial airports are vertices. In this work, we focus instead on the properties of flight delays in the US air transportation network. We analyze flight performance data in 2010 and study the topological structure of th…
▽ More
Complex networks provide a suitable framework to characterize air traffic. Previous works described the world air transport network as a graph where direct flights are edges and commercial airports are vertices. In this work, we focus instead on the properties of flight delays in the US air transportation network. We analyze flight performance data in 2010 and study the topological structure of the network as well as the aircraft rotation. The properties of flight delays, including the distribution of total delays, the dependence on the day of the week and the hour-by-hour evolution within each day, are characterized paying special attention to flights accumulating delays longer than 12 hours. We find that the distributions are robust to changes in takeoff or landing operations, different moments of the year or even different airports in the contiguous states. However, airports in remote areas (Hawaii, Alaska, Puerto Rico) can show peculiar distributions biased toward long delays. Additionally, we show that long delayed flights have an important dependence on the destination airport.
△ Less
Submitted 2 August, 2013; v1 submitted 9 April, 2013;
originally announced April 2013.
-
Dynamics in online social networks
Authors:
Przemyslaw A. Grabowicz,
Jose J. Ramasco,
Victor M. Eguiluz
Abstract:
An increasing number of today's social interactions occurs using online social media as communication channels. Some online social networks have become extremely popular in the last decade. They differ among themselves in the character of the service they provide to online users. For instance, Facebook can be seen mainly as a platform for keeping in touch with close friends and relatives, Twitter…
▽ More
An increasing number of today's social interactions occurs using online social media as communication channels. Some online social networks have become extremely popular in the last decade. They differ among themselves in the character of the service they provide to online users. For instance, Facebook can be seen mainly as a platform for keeping in touch with close friends and relatives, Twitter is used to propagate and receive news, LinkedIn facilitates the maintenance of professional contacts, Flickr gathers amateurs and professionals of photography, etc. Albeit different, all these online platforms share an ingredient that pervades all their applications. There exists an underlying social network that allows their users to keep in touch with each other and helps to engage them in common activities or interactions leading to a better fulfillment of the service's purposes. This is the reason why these platforms share a good number of functionalities, e.g., personal communication channels, broadcasted status updates, easy one-step information sharing, news feeds exposing broadcasted content, etc. As a result, online social networks are an interesting field to study an online social behavior that seems to be generic among the different online services. Since at the bottom of these services lays a network of declared relations and the basic interactions in these platforms tend to be pairwise, a natural methodology for studying these systems is provided by network science. In this chapter we describe some of the results of research studies on the structure, dynamics and social activity in online social networks. We present them in the interdisciplinary context of network science, sociological studies and computer science.
△ Less
Submitted 2 October, 2012;
originally announced October 2012.
-
Social and strategic imitation: the way to consensus
Authors:
Daniele Vilone,
José J. Ramasco,
Angel Sánchez,
Maxi San Miguel
Abstract:
Humans do not always make rational choices, a fact that experimental economics is putting on solid grounds. The social context plays an important role in determining our actions, and often we imitate friends or acquaintances without any strategic consideration. We explore here the interplay between strategic and social imitative behaviors in a coordination problem on a social network. We observe t…
▽ More
Humans do not always make rational choices, a fact that experimental economics is putting on solid grounds. The social context plays an important role in determining our actions, and often we imitate friends or acquaintances without any strategic consideration. We explore here the interplay between strategic and social imitative behaviors in a coordination problem on a social network. We observe that for interactions in 1D and 2D lattices any amount of social imitation prevents the freezing of the network in domains with different conventions, thus leading to global consensus. For interactions in complex networks, the interplay of social and strategic imitation also drives the system towards global consensus while neither dynamics alone does. We find an optimum value for the combination of imitative behaviors to reach consensus in a minimum time, and two different dynamical regimes to approach it: exponential when social imitation predominates, and power-law when strategic considerations dominate.
△ Less
Submitted 24 July, 2012; v1 submitted 23 July, 2012;
originally announced July 2012.
-
Influence of opinion dynamics on the evolution of games
Authors:
Floriana Gargiulo,
Jose J. Ramasco
Abstract:
Under certain circumstances such as lack of information or bounded rationality, human players can take decisions on which strategy to choose in a game on the basis of simple opinions. These opinions can be modified after each round by observing own or others payoff results but can be also modified after interchanging impressions with other players. In this way, the update of the strategies can bec…
▽ More
Under certain circumstances such as lack of information or bounded rationality, human players can take decisions on which strategy to choose in a game on the basis of simple opinions. These opinions can be modified after each round by observing own or others payoff results but can be also modified after interchanging impressions with other players. In this way, the update of the strategies can become a question that goes beyond simple evolutionary rules based on fitness and become a social issue. In this work, we explore this scenario by coupling a game with an opinion dynamics model. The opinion is represented by a continuous variable that corresponds to the certainty of the agents respect to which strategy is best. The opinions transform into actions by making the selection of an strategy a stochastic event with a probability regulated by the opinion. A certain regard for the previous round payoff is included but the main update rules of the opinion are given by a model inspired in social interchanges. We find that the dynamics fixed points of the coupled model is different from those of the evolutionary game or the opinion models alone. Furthermore, new features emerge such as the resilience of the fraction of cooperators to the topology of the social interaction network or to the presence of a small fraction of extremist players.
△ Less
Submitted 21 July, 2012; v1 submitted 16 July, 2012;
originally announced July 2012.
-
Dynamical Classes of Collective Attention in Twitter
Authors:
Janette Lehmann,
Bruno Gonçalves,
José J. Ramasco,
Ciro Cattuto
Abstract:
Micro-blogging systems such as Twitter expose digital traces of social discourse with an unprecedented degree of resolution of individual behaviors. They offer an opportunity to investigate how a large-scale social system responds to exogenous or endogenous stimuli, and to disentangle the temporal, spatial and topical aspects of users' activity. Here we focus on spikes of collective attention in T…
▽ More
Micro-blogging systems such as Twitter expose digital traces of social discourse with an unprecedented degree of resolution of individual behaviors. They offer an opportunity to investigate how a large-scale social system responds to exogenous or endogenous stimuli, and to disentangle the temporal, spatial and topical aspects of users' activity. Here we focus on spikes of collective attention in Twitter, and specifically on peaks in the popularity of hashtags. Users employ hashtags as a form of social annotation, to define a shared context for a specific event, topic, or meme. We analyze a large-scale record of Twitter activity and find that the evolution of hastag popularity over time defines discrete classes of hashtags. We link these dynamical classes to the events the hashtags represent and use text mining techniques to provide a semantic characterization of the hastag classes. Moreover, we track the propagation of hashtags in the Twitter social network and find that epidemic spreading plays a minor role in hastag popularity, which is mostly driven by exogenous factors.
△ Less
Submitted 1 March, 2012; v1 submitted 8 November, 2011;
originally announced November 2011.
-
Social features of online networks: the strength of intermediary ties in online social media
Authors:
Przemyslaw A. Grabowicz,
Jose J. Ramasco,
Esteban Moro,
Josep Pujol,
Victor M. Eguiluz
Abstract:
An increasing fraction of today social interactions occur using online social media as communication channels. Recent worldwide events, such as social movements in Spain or revolts in the Middle East, highlight their capacity to boost people coordination. Online networks display in general a rich internal structure where users can choose among different types and intensity of interactions. Despite…
▽ More
An increasing fraction of today social interactions occur using online social media as communication channels. Recent worldwide events, such as social movements in Spain or revolts in the Middle East, highlight their capacity to boost people coordination. Online networks display in general a rich internal structure where users can choose among different types and intensity of interactions. Despite of this, there are still open questions regarding the social value of online interactions. For example, the existence of users with millions of online friends sheds doubts on the relevance of these relations. In this work, we focus on Twitter, one of the most popular online social networks, and find that the network formed by the basic type of connections is organized in groups. The activity of the users conforms to the landscape determined by such groups. Furthermore, Twitter's distinction between different types of interactions allows us to establish a parallelism between online and offline social networks: personal interactions are more likely to occur on internal links to the groups (the weakness of strong ties), events transmitting new information go preferentially through links connecting different groups (the strength of weak ties) or even more through links connecting to users belonging to several groups that act as brokers (the strength of intermediary ties).
△ Less
Submitted 8 February, 2012; v1 submitted 20 July, 2011;
originally announced July 2011.
-
Finding statistically significant communities in networks
Authors:
Andrea Lancichinetti,
Filippo Radicchi,
Jose' Javier Ramasco,
Santo Fortunato
Abstract:
Community structure is one of the main structural features of networks, revealing both their internal organization and the similarity of their elementary units. Despite the large variety of methods proposed to detect communities in graphs, there is a big need for multi-purpose techniques, able to handle different types of datasets and the subtleties of community structure. In this paper we present…
▽ More
Community structure is one of the main structural features of networks, revealing both their internal organization and the similarity of their elementary units. Despite the large variety of methods proposed to detect communities in graphs, there is a big need for multi-purpose techniques, able to handle different types of datasets and the subtleties of community structure. In this paper we present OSLOM (Order Statistics Local Optimization Method), the first method capable to detect clusters in networks accounting for edge directions, edge weights, overlapping communities, hierarchies and community dynamics. It is based on the local optimization of a fitness function expressing the statistical significance of clusters with respect to random fluctuations, which is estimated with tools of Extreme and Order Statistics. OSLOM can be used alone or as a refinement procedure of partitions/covers delivered by other techniques. We have also implemented sequential algorithms combining OSLOM with other fast techniques, so that the community structure of very large networks can be uncovered. Our method has a comparable performance as the best existing algorithms on artificial benchmark graphs. Several applications on real networks are shown as well. OSLOM is implemented in a freely available software (http://www.oslom.org), and we believe it will be a valuable tool in the analysis of networks.
△ Less
Submitted 4 May, 2011; v1 submitted 10 December, 2010;
originally announced December 2010.
-
Information filtering in complex weighted networks
Authors:
Filippo Radicchi,
José J. Ramasco,
Santo Fortunato
Abstract:
Many systems in nature, society and technology can be described as networks, where the vertices are the system's elements and edges between vertices indicate the interactions between the corresponding elements. Edges may be weighted if the interaction strength is measurable. However, the full network information is often redundant because tools and techniques from network analysis do not work or b…
▽ More
Many systems in nature, society and technology can be described as networks, where the vertices are the system's elements and edges between vertices indicate the interactions between the corresponding elements. Edges may be weighted if the interaction strength is measurable. However, the full network information is often redundant because tools and techniques from network analysis do not work or become very inefficient if the network is too dense and some weights may just reflect measurement errors, and shall be discarded. Moreover, since weight distributions in many complex weighted networks are broad, most of the weight is concentrated among a small fraction of all edges. It is then crucial to properly detect relevant edges. Simple thresholding would leave only the largest weights, disrupting the multiscale structure of the system, which is at the basis of the structure of complex networks, and ought to be kept. In this paper we propose a weight filtering technique based on a global null model (GloSS filter), keeping both the weight distribution and the full topological structure of the network. The method correctly quantifies the statistical significance of weights assigned independently to the edges from a given distribution. Applications to real networks reveal that the GloSS filter is indeed able to identify relevantconnections between vertices.
△ Less
Submitted 15 April, 2011; v1 submitted 15 September, 2010;
originally announced September 2010.
-
Optimization of transport protocols with path-length constraints in complex networks
Authors:
Jose J. Ramasco,
Marta S. de la Lama,
Eduardo Lopez,
Stefan Boettcher
Abstract:
We propose a protocol optimization technique that is applicable to both weighted or unweighted graphs. Our aim is to explore by how much a small variation around the Shortest Path or Optimal Path protocols can enhance protocol performance. Such an optimization strategy can be necessary because even though some protocols can achieve very high traffic tolerance levels, this is commonly done by enlar…
▽ More
We propose a protocol optimization technique that is applicable to both weighted or unweighted graphs. Our aim is to explore by how much a small variation around the Shortest Path or Optimal Path protocols can enhance protocol performance. Such an optimization strategy can be necessary because even though some protocols can achieve very high traffic tolerance levels, this is commonly done by enlarging the path-lengths, which may jeopardize scalability. We use ideas borrowed from Extremal Optimization to guide our algorithm, which proves to be an effective technique. Our method exploits the degeneracy of the paths or their close-weight alternatives, which significantly improves the scalability of the protocols in comparison to Shortest Paths or Optimal Paths protocols, keeping at the same time almost intact the length or weight of the paths. This characteristic ensures that the optimized routing protocols are composed of paths that are quick to traverse, avoiding negative effects in data communication due to path-length increases that can become specially relevant when information losses are present.
△ Less
Submitted 3 June, 2010;
originally announced June 2010.
-
Agents, Bookmarks and Clicks: A topical model of Web traffic
Authors:
Mark Meiss,
Bruno Gonçalves,
José J. Ramasco,
Alessandro Flammini,
Filippo Menczer
Abstract:
Analysis of aggregate and individual Web traffic has shown that PageRank is a poor model of how people navigate the Web. Using the empirical traffic patterns generated by a thousand users, we characterize several properties of Web traffic that cannot be reproduced by Markovian models. We examine both aggregate statistics capturing collective behavior, such as page and link traffic, and individual…
▽ More
Analysis of aggregate and individual Web traffic has shown that PageRank is a poor model of how people navigate the Web. Using the empirical traffic patterns generated by a thousand users, we characterize several properties of Web traffic that cannot be reproduced by Markovian models. We examine both aggregate statistics capturing collective behavior, such as page and link traffic, and individual statistics, such as entropy and session size. No model currently explains all of these empirical observations simultaneously. We show that all of these traffic patterns can be explained by an agent-based model that takes into account several realistic browsing behaviors. First, agents maintain individual lists of bookmarks (a non-Markovian memory mechanism) that are used as teleportation targets. Second, agents can retreat along visited links, a branching mechanism that also allows us to reproduce behaviors such as the use of a back button and tabbed browsing. Finally, agents are sustained by visiting novel pages of topical interest, with adjacent pages being more topically related to each other than distant ones. This modulates the probability that an agent continues to browse or starts a new session, allowing us to recreate heterogeneous session lengths. The resulting model is capable of reproducing the collective and individual behaviors we observe in the empirical data, reconciling the narrowly focused browsing patterns of individual users with the extreme heterogeneity of aggregate traffic measurements. This result allows us to identify a few salient features that are necessary and sufficient to interpret the browsing patterns observed in our data. In addition to the descriptive and explanatory power of such a model, our results may lead the way to more sophisticated, realistic, and effective ranking and crawling algorithms.
△ Less
Submitted 27 March, 2010;
originally announced March 2010.
-
What's in a Session: Tracking Individual Behavior on the Web
Authors:
Mark Meiss,
John Duncan,
Bruno Gonçalves,
José J. Ramasco,
Filippo Menczer
Abstract:
We examine the properties of all HTTP requests generated by a thousand undergraduates over a span of two months. Preserving user identity in the data set allows us to discover novel properties of Web traffic that directly affect models of hypertext navigation. We find that the popularity of Web sites -- the number of users who contribute to their traffic -- lacks any intrinsic mean and may be…
▽ More
We examine the properties of all HTTP requests generated by a thousand undergraduates over a span of two months. Preserving user identity in the data set allows us to discover novel properties of Web traffic that directly affect models of hypertext navigation. We find that the popularity of Web sites -- the number of users who contribute to their traffic -- lacks any intrinsic mean and may be unbounded. Further, many aspects of the browsing behavior of individual users can be approximated by log-normal distributions even though their aggregate behavior is scale-free. Finally, we show that users' click streams cannot be cleanly segmented into sessions using timeouts, affecting any attempt to model hypertext navigation using statistics of individual sessions. We propose a strictly logical definition of sessions based on browsing activity as revealed by referrer URLs; a user may have several active sessions in their click stream at any one time. We demonstrate that applying a timeout to these logical sessions affects their statistics to a lesser extent than a purely timeout-based mechanism.
△ Less
Submitted 27 March, 2010;
originally announced March 2010.
-
Remembering what we like: Toward an agent-based model of Web traffic
Authors:
Bruno Goncalves,
Mark R. Meiss,
Jose J. Ramasco,
Alessandro Flammini,
Filippo Menczer
Abstract:
Analysis of aggregate Web traffic has shown that PageRank is a poor model of how people actually navigate the Web. Using the empirical traffic patterns generated by a thousand users over the course of two months, we characterize the properties of Web traffic that cannot be reproduced by Markovian models, in which destinations are independent of past decisions. In particular, we show that the div…
▽ More
Analysis of aggregate Web traffic has shown that PageRank is a poor model of how people actually navigate the Web. Using the empirical traffic patterns generated by a thousand users over the course of two months, we characterize the properties of Web traffic that cannot be reproduced by Markovian models, in which destinations are independent of past decisions. In particular, we show that the diversity of sites visited by individual users is smaller and more broadly distributed than predicted by the PageRank model; that link traffic is more broadly distributed than predicted; and that the time between consecutive visits to the same site by a user is less broadly distributed than predicted. To account for these discrepancies, we introduce a more realistic navigation model in which agents maintain individual lists of bookmarks that are used as teleportation targets. The model can also account for branching, a traffic property caused by browser features such as tabs and the back button. The model reproduces aggregate traffic patterns such as site popularity, while also generating more accurate predictions of diversity, link traffic, and return time distributions. This model for the first time allows us to capture the extreme heterogeneity of aggregate traffic measurements while explaining the more narrowly focused browsing patterns of individual users.
△ Less
Submitted 24 January, 2009;
originally announced January 2009.
-
Towards the characterization of individual users through Web analytics
Authors:
Bruno Goncalves,
Jose J. Ramasco
Abstract:
We perform an analysis of the way individual users navigate in the Web. We focus primarily in the temporal patterns of they return to a given page. The return probability as a function of time as well as the distribution of time intervals between consecutive visits are measured and found to be independent of the level of activity of single users. The results indicate a rich variety of individual…
▽ More
We perform an analysis of the way individual users navigate in the Web. We focus primarily in the temporal patterns of they return to a given page. The return probability as a function of time as well as the distribution of time intervals between consecutive visits are measured and found to be independent of the level of activity of single users. The results indicate a rich variety of individual behaviors and seem to preclude the possibility of defining a characteristic frequency for each user in his/her visits to a single site.
△ Less
Submitted 5 January, 2009;
originally announced January 2009.
-
Stability of Maximum likelihood based clustering methods: exploring the backbone of classifications (Who is keeping you in that community?)
Authors:
Muhittin Mungan,
Jose J. Ramasco
Abstract:
Components of complex systems are often classified according to the way they interact with each other. In graph theory such groups are known as clusters or communities. Many different techniques have been recently proposed to detect them, some of which involve inference methods using either Bayesian or Maximum Likelihood approaches. In this article, we study a statistical model designed for detec…
▽ More
Components of complex systems are often classified according to the way they interact with each other. In graph theory such groups are known as clusters or communities. Many different techniques have been recently proposed to detect them, some of which involve inference methods using either Bayesian or Maximum Likelihood approaches. In this article, we study a statistical model designed for detecting clusters based on connection similarity. The basic assumption of the model is that the graph was generated by a certain grouping of the nodes and an Expectation Maximization algorithm is employed to infer that grouping. We show that the method admits further development to yield a stability analysis of the groupings that quantifies the extent to which each node influences its neighbors group membership. Our approach naturally allows for the identification of the key elements responsible for the grouping and their resilience to changes in the network. Given the generality of the assumptions underlying the statistical model, such nodes are likely to play special roles in the original system. We illustrate this point by analyzing several empirical networks for which further information about the properties of the nodes is available. The search and identification of stabilizing nodes constitutes thus a novel technique to characterize the relevance of nodes in complex networks.
△ Less
Submitted 29 April, 2010; v1 submitted 8 September, 2008;
originally announced September 2008.
-
Human dynamics revealed through Web analytics
Authors:
Bruno Goncalves,
Jose J. Ramasco
Abstract:
When the World Wide Web was first conceived as a way to facilitate the sharing of scientific information at the CERN (European Center for Nuclear Research) few could have imagined the role it would come to play in the following decades. Since then, the increasing ubiquity of Internet access and the frequency with which people interact with it raise the possibility of using the Web to better obse…
▽ More
When the World Wide Web was first conceived as a way to facilitate the sharing of scientific information at the CERN (European Center for Nuclear Research) few could have imagined the role it would come to play in the following decades. Since then, the increasing ubiquity of Internet access and the frequency with which people interact with it raise the possibility of using the Web to better observe, understand, and monitor several aspects of human social behavior. Web sites with large numbers of frequently returning users are ideal for this task. If these sites belong to companies or universities, their usage patterns can furnish information about the working habits of entire populations. In this work, we analyze the properly anonymized logs detailing the access history to Emory University's Web site. Emory is a medium size university located in Atlanta, Georgia. We find interesting structure in the activity patterns of the domain and study in a systematic way the main forces behind the dynamics of the traffic. In particular, we show that both linear preferential linking and priority based queuing are essential ingredients to understand the way users navigate the Web.
△ Less
Submitted 21 May, 2008; v1 submitted 27 March, 2008;
originally announced March 2008.