Introduction to Webometrics: Quantitative Web Research for the Social Sciences
Synthesis lectures on information concepts, retrieval, and servicesPublished 1 January 2009
Mike Thelwall
Citations221
SJR quartileQ4
SJR score0.12
SNIP0.00
Generate an AI Snapshot to get a quick, structured summary of this paper.
Study Snapshot
ObjectiveStudy objective
MethodsResearch methodology
PopulationPopulation studied
Sample sizeSample sizes
OutcomesStudy outcomes here
ResultsStudy results comes here
LimitationsResearch study limitations comes here
A concise AI-generated summary of the paper will appear here once you click Generate AI Snapshot.
Abstract
Webometrics is concerned with measuring aspects of the web: web sites, web pages, parts of web pages, words in web pages, hyperlinks, web search engine results. The importance of the web itself as a c
Keywords
Computer SciencePhysics and Astronomy
ScienceEmergence of Scaling in Random Networks
36,366 Citations1999Albert-Ĺaszló Barabási, Réka Albert
A model based on these two ingredients reproduces the observed stationary scale-free distributions, which indicates that the development of large networks is governed by robust self-organizing phenomena that go beyond the particulars of the individual systems.
SIAM ReviewThe Structure and Function of Complex Networks
18,742 Citations2003Michael Newman
Developments in this field are reviewed, including such concepts as the small-world effect, degree distributions, clustering, network correlations, random graph models, models of network growth and preferential attachment, and dynamical processes taking place on networks.
Cambridge University Press eBooksSocial Network Analysis
17,041 Citations1994Stanley Wasserman, Katherine Faust
This paper presents mathematical representation of social networks in the social and behavioral sciences through the lens of dyadic and triadic interaction models, which provide insights into the structure and dynamics of relationships between actors and groups.
Computer Networks and ISDN SystemsThe anatomy of a large-scale hypertextual Web search engine
15,828 Citations1998Sergey Brin, Lawrence M. Page
This paper provides an in-depth description of Google, a prototype of a large-scale search engine which makes heavy use of the structure present in hypertext and looks at the problem of how to effectively deal with uncontrolled hypertext collections where anyone can publish anything they want.
The Content Analysis Guidebook
9,183 Citations2017Kimberly A. Neuendorf
The Content Analysis Guidebook provides an accessible core text for upper-level undergraduates and graduate students across the social sciences that unravels the complicated aspects of content analysis.
Journal of the ACMAuthoritative sources in a hyperlinked environment
9,060 Citations1999Jon Kleinberg
This work proposes and test an algorithmic formulation of the notion of authority, based on the relationship between a set of relevant authoritative pages and the set of “hub pages” that join them together in the link structure, and has connections to the eigenvectors of certain matrices associated with the link graph.
Social Science & MedicineLaboratory life: The social construction of scientific facts
2,909 Citations1982J. B. Austin
Information Processing LettersAn algorithm for drawing general undirected graphs
2,771 Citations1989Tomihisa Kamada, Satoru Kawai
Computer NetworksGraph structure in the Web
2,760 Citations2000Andrei Broder, Ravi Kumar +6 more
The study of the web as a graph yields valuable insight into web algorithms for crawling, searching and community discovery, and the sociological phenomena which characterize its evolution.
The political blogosphere and the 2004 U.S. election
2,605 Citations2005Lada A. Adamic, Natalie Glance
Differences in the behavior of liberal and conservative blogs are found, with conservative blogs linking to each other more frequently and in a denser pattern.
ACM SIGKDD Explorations NewsletterWeb usage mining
2,113 Citations2000Jaideep Srivastava, Robert Cooley +2 more
A detailed taxonomy of the work in this area, including research efforts as well as commercial offerings is provided, and a brief overview of the WebSIFT system as an example of a prototypical Web usage mining system is given.
Journal of Information ScienceUsage patterns of collaborative tagging systems
1,830 Citations2006Scott A. Golder, Bernardo A. Huberman
A dynamic model of collaborative tagging is presented that predicts regularities in user activity, tag frequencies, kinds of tags used, bursts of popularity in bookmarking and a remarkable stability in the relative proportions of tags within a given URL.
NatureAccessibility of information on the web
1,356 Citations1999Steve Lawrence, C. Lee Giles
As the web becomes a major communications medium, the data on it must be made more accessible, and search engines need to make the data more accessible.
Information science and knowledge managementCitation Analysis in Research Evaluation
1,319 Citations2005Henk F. Moed
This work focuses on assessing Basic Science Research Departments and Scientific Journals, as well as Empirical and Theoretical Chapters, and the Citation Indexes, which summarize the literature on empirical and theoretical determinants of scientific research.
Software Practice and ExperienceSoftware: Practice and Experience
1,222 Citations2006Frédéric Gervais, Benoît Fraikin
Social ForcesInvisible Colleges: Diffusion of Knowledge in Scientific Communities.
1,166 Citations1973Karen Oppenheim Mason, Diana Crane
Annual Review of Information Science and TechnologyScholarly communication and bibliometrics
1,124 Citations2002Christine L. Borgman, Jonathan Furner
The Future of Bibliometrics and its Applications: A Case Study of Agricultural Research Within the European Community Core Journals of the Rapidly Changing Research Front of "Superconductivity" is reviewed.
Elsevier eBooksCombating Web Spam with TrustRank
1,025 Citations2004Zoltán Gyöngyi, Héctor García-Molina +1 more
This paper proposes techniques to semi-automatically separate reputable, good pages from spam, and shows that they can effectively filter out spam from a significant fraction of the web, based on a good seed set of less than 200 sites.
ComputerSelf-organization and identification of Web communities
1,007 Citations2002Gary William Flake, Sandra Lawrence +2 more
This work shows that the Web self-organizes and its link structure allows efficient identification of communities and is significant because no central authority or process governs the formation and structure of hyperlinks.
Journal of DocumentationInformetric analyses on the world wide web: methodological approaches to ‘webometrics’
573 Citations1997Tomas C. Almind, Peter Ingwersen
The application of informetric methods to the World Wide Web is introduced and it is demonstrated that Denmark would seem to fall seriously behind the other Nordic countries with respect to visibility on the Net and compared to its position in scientific databases.
Journal of DocumentationThe calculation of web impact factors
466 Citations1998Peter Ingwersen
This case study reports the investigations into the feasibility and reliability of calculating impact factors for web sites, called Web Impact Factors (Web‐IF), and demonstrates that Web‐IFs are calculable with high confidence for national and sector domains whilst institutional Web‐ifs should be approached with caution.
Proceedings of the National Academy of SciencesWinners don't take all: Characterizing the competition for links on the web
461 Citations2002David M. Pennock, Gary William Flake +3 more
A simple generative model quantifies the degree to which the rich nodes grow richer, and how new (and poorly connected) nodes can compete, and accurately accounts for the true connectivity distributions of category-specific web pages, the web as a whole, and other social networks.
Journal of the American Society for Information Science and TechnologyEarlier Web usage statistics as predictors of later citation impact
396 Citations2006Tim Brody, Stevan Harnad +1 more
This paper analyses how short-term Web usage impact predicts medium-term citation impact and uses the physics e-print archive -- arXiv.org -- to test this.
NatureCitation Indexing for Studying Science
388 Citations1970Eugene Garfield
By revealing who has really influenced the course of science the Science Citation Index seems to be a valuable sociometric tool for historians and sociologists.
Online Information ReviewGoogle Scholar: the pros and the cons
387 Citations2005Péter Jacsó
There are massive content omissions presently but it feels that Google Scholar will become an excellent free tool for scholarly information discovery and retrieval with future changes in its structure.
Information Visualization: Beyond the Horizon
380 Citations2006Chaomei Chen
This paper presents a meta-modelling framework that automates the very labor-intensive and therefore time-heavy and therefore expensive and expensive process of graph drawing for knowledge domain visualization.
Journal of AdolescencePersonal information of adolescents on the Internet: A quantitative content analysis of MySpace
370 Citations2007Sameer Hinduja, Justin W. Patchin
The results indicate that the problem of personal information disclosure on MySpace may not be as widespread as many assume, and that the overwhelming majority of adolescents are responsibly using the web site.
Journal of the American Society for Information Science and TechnologyToward a basic framework for webometrics
360 Citations2004Lennart Björneborn, Peter Ingwersen
A consistent and detailed link typology and terminology is developed and a novel diagram notation is proposed to fully appreciate and investigate link structures between Web nodes in webometric analyses.
Impact of search engines on page popularity
285 Citations2004Junghoo Cho, Sourashis Roy
This paper analytically estimates how much longer it takes for a new page to attract a large number of Web users when search engines return only popular pages at the top of search results and shows that search engines can have an immensely worrisome impact on the discovery of new Web pages.
Information Processing & ManagementSearch engine coverage bias: evidence and possible causes
259 Citations2003Liwen Vaughan, Mike Thelwall
It is concluded that the coverage bias does exist but this is due not to deliberate choices of the search engines but occurs as a natural result of cumulative advantage effects of US sites on the Web.
Journal of DocumentationJOURNAL OF DOCUMENTATION
225 Citations1966
The evidence shows that, while Omeka appears to argue that adopting the Dublin Core is an integral part of Omeka ’ s mission, the platform ’ s lack of support for Dublin Core implementation makes an opposing argument.
Journal of the American Society for Information Science and TechnologyGoogle Scholar citations and Google Web/URL citations: A multi‐discipline exploratory analysis
208 Citations2007Kayvan Kousha, Mike Thelwall
Journal of the American Society for Information Science and TechnologyExtracting macroscopic information from Web links
178 Citations2001Mike Thelwall
An evaluation of Ingwersen's proposed external Web Impact Factor (WIF) for the original use of the Web: the interlinking of academic research shows that four different WIFs do, in fact, correlate with the conventional academic research measures.
Journal of the American Society for Information Science and TechnologyScientific research activity and communication measured with cybermetrics indicators
156 Citations2006Isidro F. Aguillo, Begoña Granadino +2 more
Results show that cybermetric measures could be useful for reflecting the contribution of technologically oriented institutions, increasing the visibility of developing countries, and improving the rankings based on Science Citation Index (SCI) data with known biases.
Journal of DocumentationA Tale of Two Web Spaces: Comparing Sites Using Web Impact Factors.
155 Citations1999Alastair G. Smith
For large organisations such as universities or research institutions, WIFs seem to be a useful measure of the overall influence of the web space, however for smaller spaces such as electronic journals the WIF is less reliable as a measure.
Journal of the American Society for Information Science and TechnologyBibliographic and Web citations: What is the difference?
149 Citations2003Liwen Vaughan, Debora Shaw
This work compared bibliographic and Web citations to articles in 46 journals in library and information science to find that Web citations correlated significantly with both bibliographs listed in the Social Sciences Citation Index and the ISI's Journal Impact Factor.
Journal of Broadcasting & Electronic MediaOnline Action in Campaign 2000: An Exploratory Analysis of the U.S. Political Web Sphere
147 Citations2002Kirsten Foot, Steven M. Schneider
Journal of the American Society for Information Science and TechnologyStatistical relationships between downloads and citations at the level of individual documents within a single journal
146 Citations2005Henk F. Moed
Journal of the American Society for Information ScienceInvoked on the Web
145 Citations1998Blaise Cronin, Herbert Snyder +3 more
It is argued that the Web fosters new modalities of scholarly communication and different categories of invocation are identified and analyzed in terms of their potential to inform sociometric and bibliometric analyses of academic interaction.
Journal of the American Society for Information Science and TechnologyHomophily in MySpace
138 Citations2008Mike Thelwall
Choice Reviews OnlineThe invisible Web: uncovering information sources search engines can't see
134 Citations2002
Physics Today<i>The Sociology of Science: Theoretical and Empirical Investigation</i>
130 Citations1974R. K. Merton, Dudley Shapere
Journal of DocumentationEvidence for the existence of geographic trends in university Web site interlinking
116 Citations2002Mike Thelwall
This paper develops a methodology to analyse the patterns of interlinking between university Web sites and uses it to indicate that the degree of interLinking decreases with distance, at least in the UK.
Journal of the American Society for Information Science and TechnologyWeb citation data for impact assessment: A comparison of four science disciplines
111 Citations2005Liwen Vaughan, Debora Shaw
The Web-evident impact of non-UK/USA publications might provide a balance to the geographic or cultural biases observed in ISI's data, although the stability of Web citation counts is debatable.
Online Information ReviewAn exploratory study of Google Scholar
99 Citations2007Philipp Mayr, Anne‐Kathrin Walter
The study shows deficiencies in the coverage and up‐to‐dateness of the GS index and points out which web servers are the most important data providers for this search service and which information sources are highly represented.
DOAJ (DOAJ: Directory of Open Access Journals)What is this link doing here? Beginning a fine-grained process of identifying reasons for academic hyperlink creation
95 Citations2003Thelwall Mike
Information Processing & ManagementMapping world-class universities on the web
90 Citations2008José Luís Ortega, Isidro F. Aguillo
The results show that the world-class university network is constituted from national sub-networks that merge in a central core where the principal universities of each country pull their networks toward international link relationships.
ScientometricsLinguistic patterns of academic Web use in Western Europe
87 Citations2003Mike Thelwall, Rong Tang +1 more
A survey of linguistic dimensions of Web site hosting and interlinking of the universities of sixteen European countries shows that English is the dominant language both for linking pages and for all pages, providing evidence for the multilingual character of academic use of the Web in Western Europe, at least outside the UK and Eire.
ScientometricsA new look at evidence of scholarly citation in citation indexes and from web sources
85 Citations2007Liwen Vaughan, Debora Shaw
A sample of 1,483 publications, representative of the scholarly production of LIS faculty, was searched in Web of Science, Google, and Google Scholar, and showed the potential to provide useful data for research evaluation.
Journal of the American Society for Information Science and TechnologyDo the Web sites of higher rated scholars have significantly more online impact?
84 Citations2003Mike Thelwall, Gareth Harries
It can be surmised that general Web publications are very different from scholarly journal articles and conference papers, for which scholarly quality does associate with citation impact, and that online impact should not be used to assess the quality of small groups of scholars, even within a single discipline.
ScientometricsLinks to commercial websites as a source of business information
78 Citations2004Liwen Vaughan, Guozhu Wu
Link count to a company's website was found to correlate with the company's revenue, profit, and research and development expenses, suggesting that Web hyperlinks to commercial sites can be a business performance indicator and thus a source of business information.
Journal of DocumentationInternet search engines – fluctuations in document accessibility
76 Citations2001Wouter Mettrop, Paul Nieuwenhuysen
Journal of the American Society for Information Science and TechnologyQuantitative comparisons of search engine results
75 Citations2008Mike Thelwall
Online Information ReviewGeneral patterns of tag usage among university groups in Flickr
72 Citations2008Emma Angus, Mike Thelwall +1 more
The results show that members of university image groups tend to tag in a manner that is of use to users of the system as a whole rather than merely for the tag creator.
Information Processing & ManagementA modeling approach to uncover hyperlink patterns: the case of Canadian universities
67 Citations2003Liwen Vaughan, Mike Thelwall
A multiple regression model was developed which shows that faculty quality and the language of the university are important predictors for links to a university Web site, and showed that English universities are advantaged.
Journal of the American Society for Information Science and TechnologyAssessing the impact of disciplinary research on teaching: An automatic analysis of online syllabuses
64 Citations2008Kayvan Kousha, Mike Thelwall
ScientometricsMaps of the academic web in the European Higher Education Area — an exploration of visual web indicators
60 Citations2007José Luís Ortega, Isidro F. Aguillo +2 more
The purpose is to combine methods from Social Network Analysis (SNA) and cybermetric techniques in order to ask for tendencies of integration of the European universities visible in their web presence and the role of different universities in the process of the emergence of an European Research Area.
Journal of the American Society for Information Science and TechnologyExtracting accurate and complete results from search engines: Case study windows live
59 Citations2007Mike Thelwall
This article introduces two new methods to extract extra URLs from search engines: automated query splitting and automated domain and TLD searching, and suggests that there is no way to get complete lists of matching URLs or accurate hit counts from Windows Live.
Information Processing & ManagementVisualization of the Nordic academic web: Link analysis using social network tools
58 Citations2008José Luís Ortega, Isidro F. Aguillo
Results show that the Nordic network is a cohesive network, set up by three well-defined sub-nets and it rests on the Finnish and Swedish sub-networks and it concludes that the Danish network has less visibility than other Nordic countries.
Journal of the American Society for Information Science and TechnologyEvolution, continuity, and disappearance of documents on a specific topic on the Web: A longitudinal study of “informetrics”
58 Citations2004Judit Bar‐Ilan, Bluma C. Peritz
Analysis of changes that occurred to a set of Web pages related to "informetrics" over a period of 5 years between June 1998 and June 2003 indicates that modification, disappearance, and resurfacing cannot be ignored when studying the structure and development of the Web.
Aslib ProceedingsAn initial exploration of the link relationship between UK university Web sites
54 Citations2002Mike Thelwall
The identification of groupings is encouraging evidence that Web links between universities can be mined for significant results, although it is clear that more methodological development is needed, if any but the simplest patterns are to be extracted.
Online Information ReviewBlog searching
53 Citations2007Mike Thelwall
A time series analysis of related blog postings suggests that the Danish cartoons issue attracted little attention in the English‐speaking world for four months after the initial publication, exploding only after the simultaneous start of diplomatic sanctions and a commercial boycott.
ScientometricsLocal government web sites in Finland: A geographic and webometric analysis
51 Citations2008Kim Holmberg, Mike Thelwall
It is shown that interl linking between local government bodies in Finland follows a strong geographic, or rather a geopolitical pattern and that governmental interlinking is mostly motivated by official cooperation that geographic adjacency has made possible.
Journal of the American Society for Information Science and TechnologyGraph structure in three national academic Webs: Power laws with anomalies
49 Citations2003Mike Thelwall, David Wilkinson
The graph structures of three national university publicly indexable Webs from Australia, New Zealand, and the UK were analyzed, resulting in power laws similar to those previously identified for individual university Web sites and for the AltaVista-indexed Web.
Journal of the American Society for Information Science and TechnologyVisualizing linguistic and cultural differences using Web co‐link data
48 Citations2006Liwen Vaughan
Results of the study showed that Web co-linking is not a random phenomenon and that co-link data contain useful information for Web data mining, and it is proposed that the method developed can be applied to other contexts such as analyzing relationships of different organizations or countries.
Online Information ReviewBlog search engines
48 Citations2007Mike Thelwall, Laura Carlson Hasler
Although blog searching is a useful new technique, the results are sensitive to the choice of search engine, the parameters used and the date of the search.
Journal of Information ScienceExploring the link structure of the Web with network diagrams
45 Citations2001Mike Thelwall
This article explores the network diagram as a tool to visualize the strength of the interconnection between areas of the web and reports that four different link count based weightings are possible, each highlighting a different aspect of the data.
Journal of the American Society for Information Science and TechnologyAuthor Cocitation Analysis is to intellectual structure as Web Colink Analysis is to …?
43 Citations2006Alesia Zuccala
Author Cocitation Analysis and Web Colink Analysis are examined as sister techniques in the related fields of bibliometrics and webometrics based on their data retrieval, mapping, and interpretation procedures, using mathematics as the subject in focus.
Revista española de Documentación CientíficaValoración del impacto de la información en Internet: Altavista, el “Citation Index” de la red
42 Citations1997Josep Manuel Rodríguez i Gairín
The article emphasizes the possibility to measure the impact of these pages depending on the amount of times that they are linked from external pages, similarly to the way in which the Institute for Scientific Information Citation Index works.
ScientometricsA university-centred European Union link analysis
39 Citations2008Mike Thelwall, Alesia Zuccala
The results show the expected EU dominance of the large richer Western European nations, particularly the UK and Germany, and the new EU countries are not yet integrated into the EU web but some show strong regional connections.
Journal of the American Society for Information Science and TechnologyWeb issue analysis: An integrated water resource management case study
37 Citations2006Mike Thelwall, Katie Vann +1 more
In this article Web issue analysis is introduced as a new technique to investigate an issue as reflected on the Web, a United Nations–initiated paradigm for managing water resources in an international context, particularly in developing nations.
Scientometrics'Mini small worlds' of shortest link paths crossing domain boundaries in an academic Web space
32 Citations2006Lennart Björneborn
Indicative findings suggest that personal Web page authors and computer science subsites may be importantSmall-world connectors across sites and topics in an academic Web space may counteract balkanization of the Web into insularities of disconnected and unreachable subpopulations.
Journal of the American Society for Information Science and TechnologyA statistical analysis of the web presences of European life sciences research teams
28 Citations2008Franz Barjak, Mike Thelwall
It was confirmed that research-group size and Web-presence size were important for attracting Web links, although research productivity was not, and the choice of search engine created a surprising international difference in the results, with Google perhaps giving unreliable results.
Journal of Information ScienceUK academic web links and collaboration - an exploratory study
28 Citations2007Emma Stuart, Mike Thelwall +1 more
The potential of web links to act as an indicator of collaboration through a detailed classification of 2600 links from universities to government, commercial and other domains is investigated.
Research EvaluationInvestigating triple helix relationships using URL citations: a case study of the UK West Midlands automobile industry
27 Citations2006Emma Stuart, Mike Thelwall
The potential use of web URL citations, collected through Google's API, as weak benchmarking indicators to estimate the levels of collaboration between different organisations is explored through a case study of the automobile industry in the UK West Midlands region.
Journal of the American Society for Information Science and TechnologyIdentifying and characterizing public science‐related fears from RSS feeds
23 Citations2006Mike Thelwall, Rudy Prabowo
A semi-automatic method for identifying significant public science-related concerns from a corpus of Internet-based RSS (Really Simple Syndication) feeds is described and shown to be an improvement on a previous similar system because of the introduction of feed-based aggregation.
Search engines and their public interfaces
21 Citations2007Frank McCown, Michael L. Nelson
This work provides the first in-depth quantitative analysis of the results produced by the Google, MSN and Yahoo API and WUI interfaces, and suggests that the API indexes are not older, but they are probably smaller for Google and Yahoo.
Mapping business competitive positions using web co-link analysis
20 Citations2005Liwen Vaughan, Jianqing You
It is proposed that regular data collection and analysis based on the co-link data can be used to monitor the business competitive environment and trigger early warnings on the change of the competitive landscape.
Journal of the American Society for Information Science and TechnologyLanguage evolution and the spread of ideas on the Web: A procedure for identifying emergent hybrid word family members
16 Citations2006Mike Thelwall, Liz Price
Techniques are described and tested for identifying new words from the Web, focusing on the case when the words are related to a topic and have a hybrid form with a common sequence of letters, and show the wide potential of hybrid word family investigations in linguistics and social science.
ScientometricsNational and international university departmental Web site interlinking
15 Citations2005Xuemei Li, Mike Thelwall +2 more
Whether and how link patterns differ along country and disciplinary lines between similar disciplines and similar countries is identified and five different perspectives are identified and compared for each set of departments.
Journal of Information ScienceSite navigation and its impact on the content viewed by the virtual scholar: a deep log analysis
10 Citations2007Paul Huntington, David Nicholas +1 more
A strong association was found between form of navigation and behavioural trait and use of the online searching facility increases the visibility of material irrespective of journal and age and results in a greater use of older material and a more diverse journal use compared to other online and off-line information retrieval methods.
…
