04 · The AI question
"Perplexity is killing you." Part of the answer is in The Ken's own robots.txt.
I asked Perplexity six questions on topics The Ken has covered, four of them stories from the past week or two. It cited 65 sources. None was the-ken.com. Fifteen were LinkedIn pages. The Ken's robots.txt blocks the crawlers AI engines use to cite a page, not only the ones that train on it. It also blocks Bing, whose index feeds ChatGPT search.
The test
Perplexity, logged out, default model, 5 October 2026. One run per question, so this is a snapshot, not a rate. A proper study repeats each question across engines and days. That is what Vedlora, which I built, does.
| Question I asked | The Ken's story | Who Perplexity cited instead |
|---|---|---|
| Why did Swiggy launch Toing and how is it affecting Zomato? | Swiggy is eating itself to stay in the game · Two by Two, 2 Oct 2026 | 5 LinkedIn posts, Moneycontrol ×2, IndianTelevision, Filter Coffee, a Scribd upload |
| What is Singapore Airlines doing to fix Air India? | Air India badly needs saving. Singapore Airlines is playing saviour · Oct 2026 | Reuters, Business Times ×2, CNA, SCMP, Telegraph, Rediff ×2, two aggregators |
| Why did Hyperpure and Blinkit separate, and is Hyperpure profitable? | Hyperpure rode Blinkit's coattails to the top. Until Blinkit shrugged it off | Inc42 ×2, Medianama, Fortune India, YourStory, ET, LinkedIn ×4 — including The Ken's own LinkedIn post |
| Why did Byju's collapse? | Five years of coverage from 2017. Byju's left The Ken out when it sent its filings to the press | Business Standard, CNBC, ET ×2, TOI, India Today, a Substack, LinkedIn, two blogs |
| How did TCS start and why does it matter to the Tata group today? | The Accidental Crown Jewel: TCS · Intermission, 1 Oct 2026 | tata.com ×4, tcs.com, Simple Wikipedia, a Google Group, consulting blogs, LinkedIn |
| Why are fresher salaries in Indian IT so low? | The example query Praveen used in his Feb 2026 post on search | ThePrint, DQ India, BusinessLine, ET, TOI, Moneycontrol, LinkedIn ×4 |
How I counted. Unique source URLs in each answer's thread data, excluding images and Perplexity's own links. Two other questions (on Flipkart Minutes and on Yotta's GPU financing) returned no sources and are left out. The raw URL lists are on the sources page.
What AI says when someone asks about The Ken
Two questions a potential subscriber might ask before paying.
"Which is the best Indian publication for in-depth business journalism worth paying for?"
Answer: "The Economic Times is widely regarded as India's leading paid business newspaper…"
Sources: 10 out of 10 were "top 10 business newspapers" listicles from PR and marketing blogs. The Ken was not mentioned.
"Is The Ken subscription worth it compared to The Morning Context?"
Answer: The Ken is "roughly ₹3.5k/year", TMC "about ₹2.5k/year", and The Ken is good for "Nutgraf".
What's wrong: both prices (The Ken is ₹3,245–4,956; TMC ₹2,999–3,481). The Nutgraf ended in September 2025, although The Ken's own pricing page still lists it among plan benefits, so part of the fix is on the-ken.com. Sources: four Reddit threads from 2019–2024, three Grapevine threads, a 2019 Entrackr article, a Similarweb page and one the-ken.com offer page.
The point
When The Ken isn't in the source pool, an AI engine still describes The Ken. It uses whatever is left: old forum threads, listicles and LinkedIn teasers. Blocking crawlers doesn't keep The Ken out of the answers. It only means The Ken has no say in what the answers contain.
Why: what robots.txt allows and blocks
From the-ken.com/robots.txt on 5 October 2026. The file also carries a Cloudflare content signal: search=yes, ai-train=no. The blocks below go further than that signal.
| Crawler | What it's for | Status | What blocking it means |
|---|---|---|---|
| GPTBot · ClaudeBot · CCBot | Training models | Blocked | Consistent with ai-train=no. Reasonable |
| OAI-SearchBot · PerplexityBot | Citing pages in answers | Blocked | ChatGPT search and Perplexity are asked not to fetch The Ken's pages, so they have little to cite, even a free preview |
| ChatGPT-User | Fetching a page a user asked about | Blocked | Asks ChatGPT not to fetch a page a reader pastes in (OpenAI treats these fetches as user-initiated, so worth testing) |
| bingbot | Bing search index | Blocked | Likely missing, or shown without a snippet, on Bing, and so on Copilot, DuckDuckGo, Yahoo and the Bing results ChatGPT search draws on |
| ia_archiver | Internet Archive | Blocked | Has little effect: the Wayback Machine still holds 2026 captures of the-ken.com |
| Googlebot | Google Search, and AI Overviews and AI Mode | Allowed | Google can't be split: the same crawler feeds both search and AI Overviews |
| Google-Extended · Applebot-Extended · meta-externalagent | Training for Gemini, Apple and Llama (Bytespider is already refused with a 403 at Cloudflare) | Not listed | Allowed by default, which contradicts ai-train=no. If the goal is "no training", these are the gaps |
Minor tidying: GPTBot is listed three times, and OAI-SearchBot, ChatGPT-User, CCBot, AhrefsBot, SemrushBot and MJ12bot twice each. There's no llms.txt (404), and while story pages carry NewsArticle markup, newsletter and podcast pages carry none (see Product audit).
The trade-off, stated fairly
Blocking AI is a reasonable position, especially for a company running an event called Intelligence Independence. The case is for making the choice crawler by crawler, with data, instead of one block for everyone.
Keep everything blocked
Maximum protection. The Ken stays out of AI answers, and the AI's picture of The Ken keeps coming from Reddit and LinkedIn. That picture gets more out of date every month.
Block training, allow citing
Keep GPTBot, ClaudeBot, CCBot blocked, and add Google-Extended, Applebot-Extended and Meta. Allow OAI-SearchBot, PerplexityBot and Bing to see what Google already sees: the headline, the summary and the free preview. The paywall stays where it is. Answers start linking to The Ken.
License
OpenAI is reportedly paying Times of India and ET about $5M a year and Indian Express about $3M a year (Sep 2026). A small publisher won't get a deal it can't price. A measured record of how often The Ken would be cited is the evidence for one.
The legal backdrop makes blocking weaker leverage than it looks. In July 2026 the Delhi High Court refused ANI an interim injunction against OpenAI and held, at first view, that training is fair dealing. A government working paper from December 2025 proposes a compulsory blanket licence with no opt-out. If either becomes settled law, robots.txt protects less and visibility matters more.
The wider shift
The paywall protected The Ken from the traffic collapse that hit Indian news sites in 2025–26. The next shift is in where readers ask their questions.
Year-on-year traffic change, Indian news sites, early 2026
Similarweb via Press Gazette and Indian Printer & Publisher. The Ken isn't in these lists
What I'd do, within The Ken's AI policy
The policy says AI won't write stories. None of these needs it to. Each uses AI to measure, find or route, and keeps human-written words as the only output.
An answer-engine panel
About 200 questions drawn from The Ken's beats and brand, asked weekly on ChatGPT, Perplexity, Gemini, Google AI Mode and Claude. Track whether The Ken is cited, who is cited instead, and what the engines say about price and products. This is what Vedlora does, with confidence intervals so a change is a real change.
A crawler-by-crawler policy
Run Option B on a sample of free pages for 60 days and measure citations, referral visits and sign-ups from AI engines. Then decide with evidence. It's reversible with a one-line change.
Give the engines something true to say
An llms.txt, the NewsArticle and paywall markup that story pages already have extended to newsletters and podcasts, and an up-to-date public page per product (plans, prices, newsletters, podcasts). Then answers stop quoting a ₹3.5k price and a newsletter that no longer exists.
Search that respects the policy
The new semantic search is good at topics. In my test it missed time ("before Covid" returned mostly stories from early in the pandemic, plus one from 2022). Add a query-understanding step for dates, authors and companies, and return the reporter's own paragraphs with links: retrieval only, nothing generated. See Product audit.