Statistics S/38 · Free to cite · Updated 24 Sep 2026

Original research content statistics 2026.

In 2026, 47% of journalists want PR teams to send them more data or research, and 22% want more press releases, according to Cision's 2026 State of the Media survey of 1,899 journalists. That's 2.1 journalists asking for data for every one asking for another announcement. In the same year, the share of content marketers who publish original research fell from 49% to 37% (Orbit Media, September 2026), and the most cited academic test of AI search found that adding statistics to a page raised its visibility in generated answers by 30.6% on the paper's main metric (Aggarwal et al., KDD 2024).

Most content teams already know data is supposed to "work". What they lack is the evidence split into its three separate claims: data earns backlinks, journalists want data, and AI engines cite data. Those are different mechanisms with different sources, and blending them is how weak research gets greenlit. This page keeps them apart. Every number below comes from its original publisher, and four are computed here from those inputs. Cite freely with a link.

43 sourced numbers 15 primary sources By Milan Novotný

The four numbers to remember

47%Journalists who want more data from PR · Cision 2026
+30.6%AI visibility lift from adding statistics · Princeton / IIT Delhi 2024, computed
37%Content marketers who publish original research · Orbit Media 2026
2.1xJournalists wanting data vs press releases · computed

Original research content statistics at a glance

CategoryStatisticSource
Journalists47% of journalists want more data or research from PR, the top request; 22% want more press releasesCision, 2026
Journalists88% of journalists immediately disregard pitches that miss their beat; nearly half seldom get pitches that match their coverageMuck Rack, March 2026
Journalists72% of journalists say a quarter or fewer of the pitches they get are relevantCision, 2026
AI citationsAdding statistics lifted source visibility from 19.3 to 25.2 on the GEO paper's main metric, a 30.6% gainAggarwal et al., 2024 (computed)
AI citationsOn Perplexity, statistics addition improved the Subjective Impression score by up to 37%; keyword stuffing did 10% worse than doing nothingAggarwal et al., 2024
AI citationsEarned media made up 84% of AI citations across 25 million+ links; paid content made up 0.3%Muck Rack Generative Pulse, 2026
AI citationsBranded web mentions correlate 0.664 to 0.709 with AI visibility; Domain Rating 0.266 in ChatGPTAhrefs, 2025
AI citationsEight AI search engines answered more than 60% of 1,600 news-citation queries incorrectlyTow Center, 2025
Backlinks75% of 100,000 random blog posts had zero external links (2015, dated)BuzzSumo / Moz, 2015
BacklinksPew Research articles averaged 25.7 referring domains in one BuzzSumo sample; a separate, share-biased sample of 757,317 posts averaged 3.77 (2015, dated)BuzzSumo / Moz, 2015
Marketers37% of content marketers publish original research, down from 49% a year earlierOrbit Media, 2026
Marketers22% of research publishers report strong results vs 14% of all bloggersOrbit Media, 2026
MarketersMarketers who publish data so others can cite it report strong results 1.7 times as often as all respondents (24% vs 13.9%)Orbit Media, 2026 (computed)

Do journalists want original data from brands?

Journalists want original data more than any other resource PR can send: 47% of the 1,899 journalists in Cision's 2026 State of the Media survey asked for more data or research, ahead of embargoes (45%) and expert access (42%). Muck Rack's separate 2026 survey of 897 journalists ranks original research or data third among the most valuable parts of a pitch, behind beat relevance and access to sources.

StatisticSource
47% of journalists want PR professionals to send more data or research, the most requested resource. The survey ran in January and February 2026 across 19 markets.Cision State of the Media, May 2026
45% want more embargoed or early-access information and 42% want more access to experts or interviews.Cision, 2026
22% want more press releases, 32% want more multimedia assets and 19% want more spokesperson quotes.Cision, 2026
33% say credible data or research makes them more likely to engage with a pitch. Relevance to their beat leads at 79%, a timely angle follows at 35%.Cision, 2026
66% of journalists rely on PR-provided content such as press releases, pitches and media kits for story ideas.Cision, 2026
72% say a quarter or fewer of the pitches they receive are relevant.Cision, 2026
Muck Rack's journalists ranked the most valuable pitch elements in this order: beat relevance, access to relevant sources, original research or data, high-resolution images. The survey had 897 valid responses, fielded 30 January to 2 March 2026.Muck Rack State of Journalism 2026 press release, March 2026
86% of journalists say PR pitches inspire at least some of their stories.Muck Rack, March 2026
Nearly half of journalists say they seldom receive pitches that match their coverage, and 88% immediately disregard pitches that miss their beat.Muck Rack, March 2026
78% of journalists say a pitch feels genuinely relevant when it directly affects the community their audience belongs to.Muck Rack, March 2026
Milan's read

Muck Rack's full report sits behind a form at muckrack.com/resources/research/state-of-journalism, and its press release ranks the pitch elements without percentages, so I quote the ranking. Cision and Muck Rack asked different questions of different panels, so don't combine them.

Milan's read

The number that matters sits next to the headline: 72% of journalists say a quarter or fewer of their pitches are relevant. A dataset makes a relevant pitch hard to ignore, and it does nothing for a pitch sent to the wrong beat. When I read these two surveys together, the order is clear: match the beat first, then hand over a number nobody else has, then offer the person who made it.

AI engines are the newer audience for data, and the evidence there comes from one paper that everybody quotes and few people have read.

Does adding statistics make content more visible in AI answers?

Adding statistics to a web page raised its visibility in AI-generated answers in the Princeton University and IIT Delhi paper "GEO: Generative Engine Optimization", from 19.3 to 25.2 on the paper's Position-Adjusted Word Count metric, a 30.6% gain. On the live Perplexity engine the same method's gain ranged from 8.7% to 37.2% depending on the metric.

Aggarwal and five co-authors first posted the paper to arXiv on 16 November 2023; the current version is v3 from 28 June 2024, published at KDD 2024. It built GEO-bench, a set of 10,000 queries from nine sources split into 8,000 training, 1,000 validation and 1,000 test queries, with each query paired with the top five Google results. The main test used the authors' own simulated generative engine: it fetched the top five Google results for each query and had GPT-3.5 turbo write a cited answer, run on the 1,000-query test split and averaged over five random seeds. Position-Adjusted Word Count measures how many words of that answer are attributed to your source, with citations near the top weighted more. Subjective Impression is a model-scored rating of how relevant, influential and prominent the cited source looks to a reader, rescaled so both metrics start at 19.3.

Method (GEO paper, Table 1)Position-Adjusted Word CountSubjective ImpressionChange vs no optimization (word count)
No optimization19.319.3baseline
Keyword Stuffing17.720.2-8.3%
Unique Words20.520.4+6.2%
Authoritative21.322.9+10.4%
Easy-to-Understand22.020.5+14.0%
Technical Terms22.721.4+17.6%
Fluency Optimization24.721.9+28.0%
Cite Sources24.621.9+27.5%
Statistics Addition25.223.7+30.6%
Quotation Addition27.224.7+40.9%

Source: Aggarwal et al., arXiv 2311.09735 v3, June 2024, Table 1; the right-hand column is computed here from the table. The paper states its best methods improve on baseline by 41% and 28% on the two metrics.

StatisticSource
The paper tested 9 methods: Authoritative, Statistics Addition, Keyword Stuffing, Cite Sources, Quotation Addition, Easy-to-Understand, Fluency Optimization, Unique Words and Technical Terms.Aggarwal et al., 2024 (dated)
The headline claim is that GEO "can boost visibility by up to 40%" in generative engine responses.Aggarwal et al., 2024 (dated)
On Perplexity.ai, tested on 200 queries from the GEO-bench test split, quotation addition improved Position-Adjusted Word Count by 22%, statistics addition improved Subjective Impression by up to 37%, and keyword stuffing performed 10% worse than the baseline.Aggarwal et al., 2024, Table 5 (dated)
When every source in an answer was optimized at once, statistics addition lifted the fifth-ranked source by 97.9% and cut the first-ranked source by 20.6%.Aggarwal et al., 2024, Table 2 (dated)
Cite Sources lifted the fifth-ranked source by 115.1% and cut the top-ranked source by 30.3%.Aggarwal et al., 2024, Table 2 (dated)
Milan's read

Three caveats travel with every quote of this paper. The answer engine in the main test was GPT-3.5 turbo in 2023, not today's ChatGPT, Gemini or AI Overviews. The "statistics" were added by a model rewriting the page, so the test measures how a number is presented, not whether it's true or original. And the Perplexity rerun is smaller, 200 queries with the source text uploaded as files. Its numbers also drift between table and text: the table gives 29.1 against 24.1 for quotation addition, which works out to 20.7%, while the text says 22%, and keyword stuffing's "10% worse" is 9.1% in the table.

Milan's read

I'd put the paper on a slide as "statistics addition: +31% in the lab, +9% to +37% on Perplexity", not "+40%". The 40% belongs to quotations. The more useful finding is Table 2: when everyone adds statistics, the fifth-ranked page gains 98% and the top page loses 21%. For a small site, data is one of the few levers that doesn't depend on backlinks. For the rest of the analysis, see the generative engine optimization statistics page.

The paper tells you how to write a page. It doesn't tell you which sites AI engines trust in the first place, and that's where earned media comes back in.

Which sources do AI engines cite most?

AI engines cite earned media far more than brand-owned pages: Muck Rack's Generative Pulse study found earned media made up 84% of citations across more than 25 million links from ChatGPT, Claude and Gemini in May 2026, and paid content made up 0.3%. A University of Toronto team found the same tilt, with ChatGPT drawing 93.5% of its sources on well-known brands from earned media.

StatisticSource
Earned media accounted for 84% of all AI citations across 25 million+ links from ChatGPT, Claude and Gemini, covering 17 industries.Muck Rack Generative Pulse, May 2026
Journalism made up 27% of cited sources. Across three editions since July 2025, earned media ranged from 82% to 89% and journalism from 25% to 27%.Muck Rack Generative Pulse, 2026
Paid and advertorial content made up 0.3% of AI citations.Muck Rack Generative Pulse, 2026
More than 50% of journalism citations came from articles published in the previous 12 months, and citation volume dropped sharply after six months.Muck Rack Generative Pulse press release via Yahoo Finance, May 2026
Press releases were cited 3.5 times more often in industry-trend answers than in best-of answers.Muck Rack Generative Pulse, 2026
For queries about well-known brands, ChatGPT took 93.5% of sources from earned media and 6.5% from brand sites; Claude took 87.3% from earned media; Perplexity drew 23.8% from social platforms.Chen, Wang, Chen and Koudas, University of Toronto, arXiv, September 2025
Across 75,000 brands, branded web mentions correlated 0.664 to 0.709 with visibility in ChatGPT, AI Mode and AI Overviews, against 0.266 for Domain Rating in ChatGPT.Ahrefs, December 2025
YouTube mentions showed the strongest correlation with AI visibility in the same Ahrefs study, at about 0.737.Ahrefs, 2025
Milan's read

Correlation is doing a lot of work in the Ahrefs rows, and Ahrefs says so. Brands that get mentioned a lot also tend to be big, and big brands get cited. What the Muck Rack and Toronto numbers add is direction: the engines pull from third parties, so a brand gets into answers mostly by being quoted somewhere else.

Milan's read

This is where the three claims join up. Original data gets you a journalist's article; that article becomes one of the 84% of citations that are earned; your brand rides inside it. The recency row matters too: half of journalism citations are under a year old, so a study you published in 2023 is fading from answers now. I'd plan research on a yearly cycle, the way content decay statistics suggest planning refreshes.

Citation volume says nothing about citation accuracy, and the accuracy data is sobering.

How accurately do AI engines cite original data?

AI engines misattribute sources often: the Tow Center for Digital Journalism found eight AI search engines gave incorrect answers to more than 60% of 1,600 queries asking them to identify a news article's source in March 2025. Stanford researchers found in 2023 that only 51.5% of sentences from four generative search engines were fully supported by their citations.

StatisticSource
Eight AI search engines, tested on 1,600 queries (20 publishers, 10 articles each), answered more than 60% incorrectly. Perplexity got 37% wrong; Grok 3 got 94% wrong.Tow Center for Digital Journalism, Columbia Journalism Review, March 2025
Only 51.5% of generated sentences were fully supported by citations, and 74.5% of citations supported the sentence they were attached to, across Bing Chat, NeevaAI, Perplexity and YouChat.Liu, Zhang and Liang, Stanford University, arXiv, 2023 (dated)
Milan's read

Both studies test attribution, not whether engines prefer data. They matter here for one practical reason: if your number is cited without your name, you get the influence and none of the traffic.

Milan's read

I write every statistic on my pages as a complete sentence with the number, the year and the source in it, and this is why. Engines that got attribution wrong on more than 60% of queries in the Tow test are more likely to keep your name attached when the sentence can't be separated from it. It's cheap insurance and it also reads better for humans.

If the demand side is this strong, the obvious question is why fewer marketers are supplying it.

How many content marketers publish original research?

37% of content marketers published original research in 2026, down from 49% in 2025, according to Orbit Media's survey of 1,042 content marketers. The retreat came in the same year research publishers reported strong results at 22% against a 14% benchmark.

StatisticSource
37% of content marketers publish original research in 2026, down from 49% in 2025, the steepest one-year drop in Orbit's series since 2018.Orbit Media, September 2026
22% of marketers who publish original research report strong results, 1.5 times the 14% benchmark for all respondents.Orbit Media, 2026
27% of content marketers publish original data specifically so others can cite it, and that group reports strong results at 24%.Orbit Media, 2026
13.9% of bloggers report strong results, an all-time low. 92% use AI, and AI use doesn't correlate with strong results.Orbit Media, 2026
96% of B2B marketers create thought leadership content, in a survey of 1,015 B2B marketers fielded June to August 2025.Content Marketing Institute / MarketingProfs, October 2025
Milan's read

The Orbit drop is 12 points, or a 24.5% relative fall in one year. In the same survey, the 27% who publish data specifically so others can cite it report strong results at 24%, against 13.9% for all respondents: 1.7 times the benchmark. Two separate facts sit side by side here and I won't subtract them: 47% of journalists in Cision's global panel want more data, and 37% of marketers in Orbit's panel publish it. Different populations, different questions, same direction.

Milan's read

Orbit's finding is the most counterintuitive number on this page and I believe it. AI made drafting cheap, so teams shipped more drafts and fewer studies, and a study is the one asset a model can't write for you. The content marketing ROI statistics page shows the same pattern: output up, results down.

The math behind the four computed figures comes next, so you can check it.

How we calculated the original numbers

Four figures on this page don't appear in any source. Here's the arithmetic.

  1. 2.1x: journalists who want more data versus more press releases. Cision's 2026 State of the Media asked 1,899 journalists which resources they want PR to send more of. 47% chose data or research and 22% chose press releases. 47 / 22 = 2.14. Both figures come from the same question and panel.
  2. +30.6% (lab) and +8.7% to +37.2% (Perplexity): the statistics addition lift. GEO paper Table 1 gives Statistics Addition 25.2 on Position-Adjusted Word Count and 23.7 on Subjective Impression, against 19.3 for no optimization on both. 25.2 / 19.3 = 1.306 and 23.7 / 19.3 = 1.228. Table 5 (Perplexity, 200 test queries) gives 26.2 against 24.1 and 33.9 against 24.7: 1.087 and 1.372. The test set is the 1,000-query GEO-bench test split for Table 1; the inputs are from 2024 and marked dated.
  3. 1.7x: strong results among marketers who publish data to be cited. Orbit Media's 2026 survey of 1,042 content marketers reports strong results at 24% for the 27% who publish original data specifically so others can cite it, and at 13.9% for all respondents. 24 / 13.9 = 1.73. Both figures come from the same survey and the same results question; the benchmark includes the group itself.
  4. 24.5%: the one-year fall in marketers publishing original research. Orbit Media reports 49% in 2025 and 37% in 2026 for the same annual question. (49 - 37) / 49 = 0.245, a relative decline of 24.5%, or 12 percentage points.

Change any input and the conclusion holds in direction, which is the test I apply before putting a computed number on a slide.

What the 2026 numbers say

Journalists are asking for data louder than PR is sending it. Data or research is the most requested resource in Cision's 2026 survey, 2.1 times as popular as press releases. The catch is relevance: a dataset pitched to the wrong beat lands in the 72% pile.

The AI case for statistics is real but smaller than the slide version. The Princeton and IIT Delhi paper measured a 30.6% lift for statistics addition on its main metric and 8.7% to 37.2% on Perplexity. The "+40%" belongs to quotations, and all of it was measured on a 2023 answer engine.

Earned media is the bridge between the two. AI engines take 84% of their citations from earned sources and 0.3% from paid ones, in Muck Rack's May 2026 count. Original data is the most reliable way I know into earned sources, and so into answers. Fewer marketers are publishing it this year, which I read as an opening.

FAQ

Does original research get more backlinks than regular content?

Pew Research articles averaged 25.7 referring domains in BuzzSumo and Moz's 2015 study, while the study's main sample of 757,317 well-shared posts averaged 3.77. The two figures come from different samples, so read the gap as direction, not a precise multiplier. The same study found 75% of random posts earned no external links, and all of it is dated.

What percentage of journalists want data from PR?

47% of journalists want PR teams to send more data or research, according to Cision's 2026 State of the Media survey of 1,899 journalists. That makes data the top requested resource, ahead of embargoed information (45%) and expert access (42%). Muck Rack's 2026 survey of 897 journalists ranks original research or data third among the most valuable pitch elements.

Do statistics help content get cited by AI?

Adding statistics raised visibility from 19.3 to 25.2 on the main metric of the Princeton and IIT Delhi GEO paper, a 30.6% gain on a 1,000-query test set. On Perplexity the gain ranged from 8.7% to 37.2% depending on the metric. The tests used GPT-3.5 turbo in 2023 and 2024.

What did the GEO paper find about quotations and statistics?

Quotation addition was the best single method at 41% on Position-Adjusted Word Count, and statistics addition reached 30.6%, according to Aggarwal et al. (KDD 2024, arXiv v3, June 2024). Keyword stuffing scored below doing nothing. Nine methods were tested in total.

How many marketers publish original research?

37% of content marketers publish original research in 2026, down from 49% in 2025, according to Orbit Media's survey of 1,042 marketers. Research publishers report strong results at 22%, against 14% for all respondents.

Where do AI engines get the sources they cite?

84% of AI citations come from earned media and 0.3% from paid content, according to Muck Rack's May 2026 analysis of more than 25 million links from ChatGPT, Claude and Gemini. Journalism alone makes up 27% of cited sources.

For more numbers on AI search and content performance, browse the full statistics hub.

Sources

  1. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, GEO: Generative Engine Optimization, arXiv 2311.09735 (v1 November 2023, v3 June 2024, KDD 2024)
  2. Cision, 2026 State of the Media Report (May 2026)
  3. Muck Rack, The State of Journalism 2026 (March 2026, gated)
  4. Muck Rack, 2026 State of Journalism press release, syndicated by Yahoo Finance (March 2026)
  5. Muck Rack Generative Pulse, Earned media still drives 84% of AI citations (May 2026)
  6. Muck Rack Generative Pulse press release, syndicated by Yahoo Finance (May 2026)
  7. Chen, Wang, Chen and Koudas, Generative Engine Optimization: How to Dominate AI Search, arXiv 2509.08919 (September 2025)
  8. Ahrefs, AI Brand Visibility Correlations (December 2025)
  9. Tow Center for Digital Journalism, AI Search Has a Citation Problem (March 2025)
  10. Liu, Zhang and Liang, Evaluating Verifiability in Generative Search Engines, arXiv 2304.09848 (April 2023, revised October 2023)
  11. Orbit Media, Blogging Statistics 2026 (September 2026)
  12. Content Marketing Institute / MarketingProfs, B2B Content and Marketing Trends: Insights for 2026 (October 2025)
  13. BuzzSumo / Moz, Content, Shares, and Links: Insights from Analyzing 1 Million Articles (September 2015)
  14. Semrush, The State of Content Marketing 2023 Global Report (2023)
  15. Backlinko, We Analyzed 912 Million Blog Posts (February 2019)

Methodology. Numbers were collected in September 2026 and checked against the original publisher's page, PDF or paper, not against other statistics roundups. The GEO paper figures were read from the paper's own tables in arXiv version 3. Muck Rack's State of Journalism report is gated, so its figures come from Muck Rack's own press release. Claims about backlinks, journalists and AI citations come from different studies and are kept in separate sections; where two surveys cover a similar question (Cision's 47% and Muck Rack's pitch ranking), both are shown and not combined. Stats older than 24 months are marked dated. Last updated 24 September 2026.

Need one of these numbers for your own piece?

Cite it with a link to this page and the original source. Every figure above is traceable to its publisher. If you spot a number that has gone stale, email me and I will fix it within a week.

Email Milan