Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The network was generated using email data from a large European research institution. For a period from October 2003 to May 2005 (18 months) we have anonymized information about all incoming and outgoing email of the research institution. For each sent or received email message we know the time, the sender and the recipient of the email. Overall we have 3,038,531 emails between 287,755 different email addresses. Note that we have a complete email graph for only 1,258 email addresses that come from the research institution. Furthermore, there are 34,203 email addresses that both sent and received email within the span of our dataset. All other email addresses are either non-existing, mistyped or spam.
Given a set of email messages, each node corresponds to an email address. We create a directed edge between nodes i and j, if i sent at least one message to j.
Enron email communication network covers all the email communication within a dataset of around half million emails. This data was originally made public, and posted to the web, by the Federal Energy Regulatory Commission during its investigation. Nodes of the network are email addresses and if an address i sent at least one email to address j, the graph contains an undirected edge from i to j. Note that non-Enron email addresses act as sinks and sources in the network as we only observe their communication with the Enron email addresses.
The Enron email data was originally released by William Cohen at CMU.
Wikipedia is a free encyclopedia written collaboratively by volunteers around the world. Each registered user has a talk page, that she and other users can edit in order to communicate and discuss updates to various articles on Wikipedia. Using the latest complete dump of Wikipedia page edit history (from January 3 2008) we extracted all user talk page changes and created a network.
The network contains all the users and discussion from the inception of Wikipedia till January 2008. Nodes in the network represent Wikipedia users and a directed edge from node i to node j represents that user i at least once edited a talk page of user j.
The dynamic face-to-face interaction networks represent the interactions that happen during discussions between a group of participants playing the Resistance game. This dataset contains networks extracted from 62 games. Each game is played by 5-8 participants and lasts between 45--60 minutes. We extract dynamically evolving networks from the free-form discussions using the ICAF algorithm. The extracted networks are used to characterize and detect group deceptive behavior using the DeceptionRank algorithm.
The networks are weighted, directed and temporal. Each node represents a participant. At each 1/3 second, a directed edge from node u to v is weighted by the probability of participant u looking at participant v or the laptop. Additionally, we also provide a binary version where an edge from u to v indicates participant u looks at participant v (or the laptop).
Stanford Network Analysis Platform (SNAP) is a general purpose, high performance system for analysis and manipulation of large networks. Graphs consists of nodes and directed/undirected/multiple edges between the graph nodes. Networks are graphs with data on nodes and/or edges of the network.
The core SNAP library is written in C++ and optimized for maximum performance and compact graph representation. It easily scales to massive networks with hundreds of millions of nodes, and billions of edges. It efficiently manipulates large graphs, calculates structural properties, generates regular and random graphs, and supports attributes on nodes and edges. Besides scalability to large graphs, an additional strength of SNAP is that nodes, edges and attributes in a graph or a network can be changed dynamically during the computation.
SNAP was originally developed by Jure Leskovec in the course of his PhD studies. The first release was made available in Nov, 2009. SNAP uses a general purpose STL (Standard Template Library)-like library GLib developed at Jozef Stefan Institute. SNAP and GLib are being actively developed and used in numerous academic and industrial projects.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
Dataset information
The network was generated using email data from a large European research
institution. For a period from October 2003 to May 2005 (18 months) we have
anonymized information about all incoming and outgoing email of the research
institution. For each sent or received email message we know the time, the
sender and the recipient of the email. Overall we have 3,038,531 emails between
287,755 different email addresses. Note that we have a complete email graph for
only 1,258 email addresses that come from the research institution.
Furthermore, there are 34,203 email addresses that both sent and received email
within the span of our dataset. All other email addresses are either
non-existing, mistyped or spam.
Given a set of email messages, each node corresponds to an email address. We
create a directed edge between nodes i and j, if i sent at least one message to
j.
Dataset statistics
Nodes 265214
Edges 420045
Nodes in largest WCC 224832 (0.848)
Edges in largest WCC 395270 (0.941)
Nodes in largest SCC 34203 (0.129)
Edges in largest SCC 151930 (0.362)
Average clustering coefficient 0.3093
Number of triangles 267313
Fraction of closed triangles 0.004106
Diameter (longest shortest path) 13
90-percentile effective diameter 4.5
Source (citation)
J. Leskovec, J. Kleinberg and C. Faloutsos. Graph Evolution: Densification and
Shrinking Diameters. ACM Transactions on Knowledge Discovery from Data (ACM
TKDD), 1(1), 2007.
Files
File Description
email-EuAll.txt.gz Email network of a large European Research Institution
Dataset information
Enron email communication network covers all the email communication within a
dataset of around half million emails. This data was originally made public,
and posted to the web, by the Federal Energy Regulatory Commission during its
investigation. Nodes of the network are email addresses and if an address i
sent at least one email to address j, the graph contains a directed edge from i
to j. Note that non-Enron email addresses act as sinks and sources in the
network as we only observe their communication with the Enron email addresses.
The Enron email data was originally released by William Cohen at CMU.
Dataset statistics
Nodes 36692
Edges 367662
Nodes in largest WCC 33696 (0.918)
Edges in largest WCC 361622 (0.984)
Nodes in largest...
Facebook
Twitterhttps://redu.unicamp.br/api/datasets/:persistentId/versions/1.0/customlicense?persistentId=doi:10.25824/redu/CAVFDThttps://redu.unicamp.br/api/datasets/:persistentId/versions/1.0/customlicense?persistentId=doi:10.25824/redu/CAVFDT
This dataset contains complementary data to the paper "A Hybrid Matheuristic for the Spread of Influence on Social Networks" [1], which proposes a matheuristic for combinatorial optimization problems involving the spread of information in social networks. For the computational experiments discussed in that paper, we provide: - Two sets of instances, originally obtained from [2-6]; - The solutions attained by exact and heuristic methods; - The collected results; - The matheuristic source code; The directories "benchmark_*/instances/" contain files that describe the sets of instances. Each instance is associated with a graph containing {n} vertices and {m} edges. The first {m} lines of each file contain: {u} {v} where {u} and {v} identify a pair of vertices that determines an undirected edge. The next line contains {n} integers corresponding to the costs of the vertices. The last line contains {n} integers corresponding to the thresholds of the vertices. The directories "benchmark_*/solutions_*/" contain files describing feasible solutions for the corresponding sets of instances. The first line of each file contains: {s} where {s} is the number of vertices in the target set. Each of the next {s} lines contains: {v} where {v} identifies a target. The last line contains an integer that represents the target set cost. The directory "hmf_source_code/" contains an implementation of the matheuristic framework proposed in [1], namely, HMF. This work was supported by grants from Santander Bank, the Brazilian National Council for Scientific and Technological Development (CNPq), the São Paulo Research Foundation (FAPESP), the Fund for Support to Teaching, Research and Outreach Activities (FAEPEX), and the Coordination for the Improvement of Higher Education Personnel (CAPES), all in Brazil. Caveat: The opinions, hypotheses and conclusions or recommendations expressed in this material are the sole responsibility of the authors and do not necessarily reflect the views of Santander, CNPq, FAPESP, FAEPEX, or CAPES. References [1] F. C. Pereira, P. J. de Rezende, and T. Yunes. A Hybrid Matheuristic for the Spread of Influence on Social Networks. 2024. Submitted. [2] S. Raghavan and R. Zhang. A branch-and-cut approach for the weighted target set selection problem on social networks. 2024. https://doi.org/10.1287/ijoo.2019.0012 [3] J. Leskovec and A. Krevl. SNAP Datasets: Stanford Large Network Dataset Collection. 2024. https://snap.stanford.edu/data [4] R. A. Rossi and N. K. Ahmed. The Network Data Repository with Interactive Graph Analytics and Visualization. 2022. https://networkrepository.com [5] J. Kunegis. KONECT – The Koblenz Network Collection. 2013. http://dl.acm.org/citation.cfm?id=2488173 [6] O. Lesser, L. Tenenboim-Chekina, L. Rokach, and Y. Elovici. Intruder or Welcome Friend: Inferring Group Membership in Online Social Networks. 2013. https://doi.org/10.1007/978-3-642-37210-0_40
Facebook
Twitterhttps://redu.unicamp.br/api/datasets/:persistentId/versions/1.1/customlicense?persistentId=doi:10.25824/redu/ZGX0H7https://redu.unicamp.br/api/datasets/:persistentId/versions/1.1/customlicense?persistentId=doi:10.25824/redu/ZGX0H7
This dataset contains complementary data to the paper "A Row Generation Algorithm for Finding Optimal Burning Sequences of Large Graphs" [1], which proposes an exact algorithm for the Graph Burning Problem, an NP-hard optimization problem that models a form of contagion diffusion on social networks. Concerning the computational experiments discussed in that paper, we make available: - Four sets of instances; - The optimal (or best known) solutions obtained; - The source code; - An Appendix with additional details about the results. The "delta" input sets include graphs that are real-world networks [1,2], while the "grid" input set contains graphs that are square grids. The directories "delta_10K_instances", "delta_100K_instances", "delta_4M_instances" and "grid_instances" contain files that describe the sets of instances. The first two lines of each file contain: {n} {m} where {n} and {m} are the number of vertices and edges in the graph. Each of the next {m} lines contains: {u} {v} where {u} and {v} identify a pair of vertices that determines an undirected edge. The directories "delta_10K_solutions", "delta_100K_solutions", "delta_4M_solutions" and "grid_solutions" contain files that describe the optimal (or best known) solutions for the corresponding sets of instances. The first line of each file contains: {s} where {s} is the number of vertices in the burning sequence. Each of the next {s} lines contains: {v} where {v} identifies a fire source. The fire sources are listed in the same order that they appear in a burning sequence of length {s}. The directory "source_code" contains the implementations of the exact algorithm proposed in the paper [1], namely, PRYM. Lastly, the file "appendix.pdf" presents additional details on the results reported in the paper. This work was supported by grants from Santander Bank, Brazil, Brazilian National Council for Scientific and Technological Development (CNPq), Brazil, São Paulo Research Foundation (FAPESP), Brazil and Fund for Support to Teaching, Research and Outreach Activities (FAEPEX). Caveat: the opinions, hypotheses and conclusions or recommendations expressed in this material are the sole responsibility of the authors and do not necessarily reflect the views of Santander, CNPq, FAPESP or FAEPEX. References [1] F. C. Pereira, P. J. de Rezende, T. Yunes and L. F. B. Morato. A Row Generation Algorithm for Finding Optimal Burning Sequences of Large Graphs. Submitted. 2024. [2] Jure Leskovec and Andrej Krevl. SNAP Datasets: Stanford Large Network Dataset Collection. 2024. https://snap.stanford.edu/data [3] Ryan A. Rossi and Nesreen K. Ahmed. The Network Data Repository with Interactive Graph Analytics and Visualization. In: AAAI, 2022. https://networkrepository.com
Facebook
TwitterApache License, v2.0https://www.apache.org/licenses/LICENSE-2.0
License information was derived automatically
This dataset contains a directed web graph of Stanford University webpages from 2002.
Each node represents a page from stanford.edu, and each directed edge represents a hyperlink from one page to another.
It is useful for tasks such as: - network analysis - graph mining - community detection - link analysis - ranking and centrality studies - large-scale graph algorithm benchmarking
| Metric | Value |
|---|---|
| Nodes | 281,903 |
| Edges | 2,312,497 |
| Nodes in largest WCC | 255,265 (0.906) |
| Edges in largest WCC | 2,234,572 (0.966) |
| Nodes in largest SCC | 150,532 (0.534) |
| Edges in largest SCC | 1,576,314 (0.682) |
| Average clustering coefficient | 0.5976 |
| Number of triangles | 11,329,473 |
| Fraction of closed triangles | 0.002889 |
| Diameter (longest shortest path) | 674 |
| 90-percentile effective diameter | 9.7 |
stanford.edu)Because the graph is directed, a link from page A to page B does not imply a link from page B to page A.
web-Stanford.txt.gz — Stanford web graph from 2002This dataset can be used for: - studying the structure of real-world web graphs - analyzing weakly and strongly connected components - computing graph centrality and prestige measures - evaluating shortest path and diameter algorithms - exploring clustering, triangles, and transitivity - benchmarking large-scale graph processing systems
J. Leskovec, K. Lang, A. Dasgupta, M. Mahoney.
Community Structure in Large Networks: Natural Cluster Sizes and the Absence of Large Well-Defined Clusters.
Internet Mathematics, 6(1): 29–123, 2009.
If you use this dataset, please cite:
J. Leskovec, K. Lang, A. Dasgupta, and M. Mahoney. Community Structure in Large Networks: Natural Cluster Sizes and the Absence of Large Well-Defined Clusters. Internet Mathematics 6(1), 29–123, 2009.
Facebook
TwitterAttribution-ShareAlike 4.0 (CC BY-SA 4.0)https://creativecommons.org/licenses/by-sa/4.0/
License information was derived automatically
This large corpus can be used to train scientific paper summarization models that utilize citations, facilitating research in supervised methods.
Previous datasets for scientific document summarization are small with only several dozen articles. This dataset includes 1000 examples which is much larger than the prior works.
I acquired this dataset from here in XML format. The CL-Scisumm project developed the first large-scale, human-annotated Scisumm dataset, ScisummNet. It provides over 1,000 papers in the ACL anthology network with their citation networks (e.g. citation sentences, citation counts) and their comprehensive, manual summaries.
The text column has every token of the research paper, and the summary column consists of summaries of the scientific paper.
This dataset is possible by the CL-Scisumm shared task, which has been organized since 2014 for papers in the computational linguistics and NLP domain.
This dataset should be trained with SOTA models and perform better than the model proposed by the SCisummNet.
Not seeing a result you expected?
Learn how you can add new datasets to our index.
Facebook
TwitterAttribution 4.0 (CC BY 4.0)https://creativecommons.org/licenses/by/4.0/
License information was derived automatically
The network was generated using email data from a large European research institution. For a period from October 2003 to May 2005 (18 months) we have anonymized information about all incoming and outgoing email of the research institution. For each sent or received email message we know the time, the sender and the recipient of the email. Overall we have 3,038,531 emails between 287,755 different email addresses. Note that we have a complete email graph for only 1,258 email addresses that come from the research institution. Furthermore, there are 34,203 email addresses that both sent and received email within the span of our dataset. All other email addresses are either non-existing, mistyped or spam.
Given a set of email messages, each node corresponds to an email address. We create a directed edge between nodes i and j, if i sent at least one message to j.
Enron email communication network covers all the email communication within a dataset of around half million emails. This data was originally made public, and posted to the web, by the Federal Energy Regulatory Commission during its investigation. Nodes of the network are email addresses and if an address i sent at least one email to address j, the graph contains an undirected edge from i to j. Note that non-Enron email addresses act as sinks and sources in the network as we only observe their communication with the Enron email addresses.
The Enron email data was originally released by William Cohen at CMU.
Wikipedia is a free encyclopedia written collaboratively by volunteers around the world. Each registered user has a talk page, that she and other users can edit in order to communicate and discuss updates to various articles on Wikipedia. Using the latest complete dump of Wikipedia page edit history (from January 3 2008) we extracted all user talk page changes and created a network.
The network contains all the users and discussion from the inception of Wikipedia till January 2008. Nodes in the network represent Wikipedia users and a directed edge from node i to node j represents that user i at least once edited a talk page of user j.
The dynamic face-to-face interaction networks represent the interactions that happen during discussions between a group of participants playing the Resistance game. This dataset contains networks extracted from 62 games. Each game is played by 5-8 participants and lasts between 45--60 minutes. We extract dynamically evolving networks from the free-form discussions using the ICAF algorithm. The extracted networks are used to characterize and detect group deceptive behavior using the DeceptionRank algorithm.
The networks are weighted, directed and temporal. Each node represents a participant. At each 1/3 second, a directed edge from node u to v is weighted by the probability of participant u looking at participant v or the laptop. Additionally, we also provide a binary version where an edge from u to v indicates participant u looks at participant v (or the laptop).
Stanford Network Analysis Platform (SNAP) is a general purpose, high performance system for analysis and manipulation of large networks. Graphs consists of nodes and directed/undirected/multiple edges between the graph nodes. Networks are graphs with data on nodes and/or edges of the network.
The core SNAP library is written in C++ and optimized for maximum performance and compact graph representation. It easily scales to massive networks with hundreds of millions of nodes, and billions of edges. It efficiently manipulates large graphs, calculates structural properties, generates regular and random graphs, and supports attributes on nodes and edges. Besides scalability to large graphs, an additional strength of SNAP is that nodes, edges and attributes in a graph or a network can be changed dynamically during the computation.
SNAP was originally developed by Jure Leskovec in the course of his PhD studies. The first release was made available in Nov, 2009. SNAP uses a general purpose STL (Standard Template Library)-like library GLib developed at Jozef Stefan Institute. SNAP and GLib are being actively developed and used in numerous academic and industrial projects.