(reading time: 10 min)
In his article ‘There Is No Such Thing as Open Source Intelligence’, Hatfield argues that OSINT is an incoherent concept that should be abandoned (Hatfield 2024: 397). I do not share that conclusion, but some of his arguments touch on a question I left open when I proposed a working definition of OSINT: what exactly makes a source ‘open’?
One of his examples illustrates the problem well. In the early 1950s, one of the best British sources on the Soviet nuclear programme was a local newspaper published near the Semipalatinsk test site in Kazakhstan. When the number of people in the region grew in the run-up to a test, the paper reported on it, and GCHQ analysts read such reports as indicators of a coming test. After a test, the pages themselves could be checked for radioactive debris. The paper was freely available in the region, yet the British had great difficulty getting hold of copies. Michael Goodman, who interviewed the former GCHQ officer involved, concluded according to Hatfield that the officer would never have called any of this OSINT (Hatfield 2024: 401).
For Hatfield, the example shows that public availability does not work as a criterion to demarcate OSINT. I read it the other way around. The newspaper was open. The British could not get to the region where it was available. These are two different facts, and only the first one says something about the newspaper. In this blogpost I therefore argue that openness is set on the source side, by whoever controls access, and is tested against a generic member of the public rather than against what a particular analyst can reach. In short: we should not mistake the restraints of the observer for the characteristics of the source.
Available to whom?
Hatfield’s main argument against availability is a list of barriers. Material is not available in any meaningful sense, he argues, if you do not speak its language, if it requires expertise you lack, if it is not indexed by a search engine, if it sits on an intranet you cannot access, if it is hosted on the dark web, if the subscription is too costly, or if it requires a membership for which you are ineligible. He adds what Omand calls PROTINT: protected information held in government or private-sector databases (Hatfield 2024: 400).
When we sort this list by who imposes the barrier, it turns out to do two different categories. Language, expertise, indexing, the dark web and the price of a subscription describe the analyst: a skill set, a training record, a budget. Change the analyst and the barrier changes. The intranet, the membership for which you are ineligible and the protected databases of PROTINT describe the source. These are access regimes designed to exclude the public, and they only change if whoever controls the source changes them.
The second group is therefore not a set of counterexamples: these sources are not open, and the criterion correctly says so. The first group tells us something about the analyst and nothing about the source.
Applied to other collection disciplines, the test leads to odd results. An intercept in a language nobody in the building reads would then not be SIGINT, and imagery that needs a trained interpreter would not be IMINT. Doctrine does not work that way: translation and decryption belong to processing and exploitation, after collection, and not to the definition of a discipline. Hatfield himself notes that translation was a central function of the Foreign Broadcast Information Service and its successors (2024: 400): language was treated as a capability to build, not as a property of the sources.
What do the definitions say?
So what do the existing definitions say on this point? None of them ask what a particular analyst can reach. They all ask about a hypothetical member of the public.
The US Intelligence Community defines open source information as publicly available material ‘that anyone can lawfully obtain by request, purchase, or observation’ (Best and Cumming 2007: 5-6). The Attorney General procedures under Executive Order 12333 add that commercial data only count as publicly available if a person or company outside the US government could acquire the same data in the same way, and not if the data are so tailored to government use that a similarly situated private purchaser could not obtain them (ODNI 2020: s. 10.17).
The Berkeley Protocol takes the same approach. Open source information is what any member of the public can observe, purchase or request without special legal status or unauthorised access. Paid services count when they are open to everyone, but not when access is limited to groups such as law enforcement or licensed investigators (OHCHR and Human Rights Center 2020: paras 14 and 17). Dutch law speaks of ‘voor een ieder toegankelijke informatiebronnen’, information sources accessible to anyone (art. 25(1)(a) Wiv 2017).
All these definitions work with a counterfactual public. The question is not whether you can obtain the material, but whether anyone could. Van Puyvelde and Tabárez Rienzi make the same point when they describe public information as accessible to all in theory (2025: 3). This also answers Hatfield’s complaint that OSINT is only demarcated by negation, a ‘junk drawer’ defined over and against the classified disciplines (2024: 399). Accessibility to a generic public is a positive property of a source, and one that can be tested.
What about an analyst without restraints?
A thought experiment helps to sharpen the distinction. Imagine an analyst at a service that operates without any legal restraints, with one exception: he may not hack. He may buy from data brokers without limit, download leaked datasets, work on the dark web, scrape websites in breach of their terms of service, and keep whatever he finds. The only thing he may not do is defeat a technical access control. Does he have more open sources than an analyst at the AIVD?
I believe he does not. He has the same open sources, but a much larger part of them is within his reach. Nothing he is allowed to do changes the access regime of a single source.
The one exception is revealing. Hacking means defeating an access control, and an access control is exactly what makes a source closed. As a first approximation, the boundary of OSINT therefore lies where the prohibition on intrusion lies; every other restraint is a restraint on the analyst.
Leaked material follows from this. Once a dataset can be downloaded by any visitor of a forum, the access control has already failed, and it failed at the source. The material is open, and I think we should say so. Whether a given analyst may use it is a separate question.
The same goes for payment. A commercial dataset sold to anyone who pays is open; the analyst with the larger budget simply has a larger reach. The Dutch oversight committee CTIVD takes a stricter view, holding that commercial data available only against payment may fall outside the sources accessible to anyone and therefore should have a different legal basis (CTIVD 2022: 16). This does not contradict the argument: the CTIVD is not defining a discipline but interpreting a power that allows services to collect without prior authorisation, and for that purpose a strict reading makes sense. The strictness belongs to the power, not to the source.
Was it ever easy?
Practitioners made this distinction long before OSINT had a name. Becker, comparing Soviet and American access to published information, distinguished between priced Soviet publications that Americans could buy and items marked as not available for export (1957: 37). The first group was open, the second was closed by the distributor. A few years later, a survey of open sources on Soviet military affairs noted that official military regulations circulated without restriction inside the Soviet Union but required special effort to procure, so that few copies reached analysts (Moore 1963: 102). The author still called them open sources. Difficulty of acquisition was a collection problem to be solved, not a reason to reclassify the source.
The reverse also occurs. In my research on the early history of OSINT I concluded that the Venetian avvisi probably do not qualify, because they circulated among trusted contacts rather than the public, and obtaining them through agents looks more like HUMINT (Block 2024: 98). That material was hard to get and not open, which are two separate findings. Open sources can also be acquired covertly. The Chinese Communist Party built the largest collection of foreign publications in China partly through a covert purchasing post in Hong Kong (Jiang and Minami 2023: 46), and the German Sektion IIIb did something similar during the First World War (Block 2024: 102).
Open source and overt collection are therefore not the same, and once that is clear the Semipalatinsk case is no longer a puzzle. A newspaper that is freely available in the region is an open source. The British problem was access to a denied area, a hard but entirely ordinary collection problem. Testing the paper for radioactive traces is a measurement on a physical sample, and calling that MASINT supports a taxonomy based on how information is collected rather than undermining it. Reading local news for indicators of a coming test is what exploiting open sources has always been about.
The same confusion appears in a contemporary form. A widely read vendor guide states that information requiring specialist skills, tools or techniques to access cannot reasonably be considered open source (Borges 2024). Taken literally, that would exclude most of what Bellingcat does. The intended target is probably intrusion rather than skill, but the wording shows how easily the competence of the analyst ends up in the definition of the source.
Which brings me to what I find the most useful part of Hatfield’s list. Languages, searching beyond the indexed web, working safely on the dark web, acquiring hard-to-get publications and the domain knowledge to recognise what matters: that is OSINT tradecraft, more or less the syllabus of any serious OSINT course. If open meant easy, there would be no tradecraft to teach, and Hatfield would be right that the label adds nothing.
What remains open?
First, openness is relational and bounded. ‘The public’ is always a public in a certain place at a certain time. Becker’s export-restricted titles were open in Moscow and closed in Washington, and the Semipalatinsk paper was available in Kazakhstan but not in London. A critic could say that openness therefore does depend on who is asking. My answer is that in these cases the boundary was drawn by whoever controlled distribution, not by the abilities of the reader, which keeps it on the source side.
Second, the line is contested at the margins, and the prohibition on intrusion is only a first approximation. Registration walls, moderated closed groups, pretext accounts and scraping against terms of service involve no technical intrusion, but all involve an access regime that discriminates to some degree. The Berkeley Protocol considers registration compatible with openness when registration is open to all, which seems right to me, but that does not settle the harder cases. I have not yet found a satisfying answer to these.
Third, some material is open in principle but reachable by almost nobody. Grey literature is a good example: Serscikov notes that collecting it often requires physical presence, which pushes it towards human collection (2024: 1037). One could argue that calling such material open does little practical work. I would still keep the distinction, because the set of open sources is the same for all of us, while the subset any of us can reach is not.
Finally, this blogpost does not address all of Hatfield’s arguments. His points on taxonomy and on the heterogeneity of what is collected under the OSINT label deserve a separate discussion.
Conclusion
Hatfield’s barriers do not show that availability fails as a criterion. They show that we tend to mix up two questions: whether a source is open, and whether a particular analyst can or may reach it. I believe openness is set on the source side, by whoever controls access, and tested against a generic member of the public rather than against the analyst who happens to be asking. Language, skills, budget, indexing and legal position determine what an analyst can reach, not what is open.
I am aware that the line I have drawn through leaked material will not satisfy everyone, as it makes a hacked dataset an open source the moment it is posted. I think that is the correct result, provided the legality of using it is treated as a separate question. As with my working definition, I am not fully convinced that I have all the edge cases right, so comments and suggestions are always welcome.
References
Becker, J. (1957) ‘Comparative Survey of Soviet and US Access to Published Information’, Studies in Intelligence, 1 (Fall), 35-42.
Best, R. and A. Cumming (2007) Open Source Intelligence (OSINT): Issues for Congress. CRS Report RL34270.
Block, L. (2024) ‘The long history of OSINT’, Journal of Intelligence History, 23(2), 95-109. https://doi.org/10.1080/16161262.2023.2224091
Borges, E. (2024) ‘What is open-source intelligence (OSINT)?’, Recorded Future, updated 24 June (first published 19 February 2019). https://www.recordedfuture.com/blog/open-source-intelligence-definition
CTIVD (2022) Toezichtsrapport nr. 74. Automated OSINT: tools en bronnen voor openbronnenonderzoek.
Hatfield, J.M. (2024) ‘There Is No Such Thing as Open Source Intelligence’, International Journal of Intelligence and CounterIntelligence, 37(2), 397-418. https://doi.org/10.1080/08850607.2023.2172367
Jiang, H. and K. Minami (2023) ‘The Eyes and Ears of the Dragon: Open-source intelligence and Chinese foreign policy during the Cold War’, Journal of Cold War Studies, 25(2), 41-63.
Moore, D. (1963) ‘Open Sources on Soviet Military Affairs’, Studies in Intelligence, 7 (Summer), 101-113.
ODNI (2020) Intelligence Activities Procedures Approved by the Attorney General Pursuant to Executive Order 12333, 23 December. https://www.intelligence.gov/assets/documents/702-documents/declassified/AGGs/
OHCHR and Human Rights Center, UC Berkeley (2020) Berkeley Protocol on Digital Open Source Investigations, advance version, HR/PUB/20/2.
Serscikov, G. (2024) ‘Grey literature in the intelligence domain: twilight or revival?’, Intelligence and National Security, 39(6), 1028-1050. https://doi.org/10.1080/02684527.2024.2372119
Van Puyvelde, D. and F. Tabárez Rienzi (2025) ‘The rise of open-source intelligence’, European Journal of International Security, 1-15. https://doi.org/10.1017/eis.2024.61
Wet op de inlichtingen- en veiligheidsdiensten 2017, art. 25 and 38. https://wetten.overheid.nl/BWBR0039896