Three types of reptiles that can be confirmed according to official documents:
OpenAI has released three distinct user agents that can be controlled separately by the platform: OAI-SearchBot: Used to create search indexes, ensuring the website appears in ChatGPT's search results.; ChatGPT-User: Represents the user's immediate request to read a webpage within a conversation.; GPTBot: Used for collecting training data for the model. Each user agent has its own independent configuration in the robots.txt file. Blocking GPTBot does not mean the website will be excluded from ChatGPT's search results. Websites that want to be visible in search results should at least allow OAI-SearchBot.
The three OpenAI user agents serve different purposes.| User agent | Applications | How to get ChatGPT to cite sources when searching |
|---|
| OAI-SearchBot | Create a search index and display links. | Must be approved. |
| ChatGPT-User | Users expect immediate page loading. | Recommended: Release |
| GPTBot | Gathering training data for the model | Decide independently based on the content licensing policy. |
How do citations arise? And what distinguishes between those that can be confirmed and those that cannot?
Confirmed: ChatGPT searches across multiple sources to generate answers, and provides links to the sources alongside the answers. Unconfirmed: The order of sources, the reasoning behind choosing A over B, and the weighting of content features – these are areas where OpenAI has not publicly disclosed information. Most "ChatGPT Citation Factor Studies" are conducted by third parties who make inferences based on samples. While these studies can be referenced, it's important to note that they are based on inference and should not be presented as platform rules. Our position: Focus on verifying the fundamentals and rely on observation rather than speculation.
Our observation method: Fixed query set
The method involves creating a fixed list of questions and regularly re-testing ChatGPT using these questions, recording the results. Key points include: Choose questions that reflect actual customer decision-making processes (e.g., comparing services, asking about pricing). Avoid simply testing brand names, as this only reflects existing knowledge.; Start a new conversation each time to avoid context influencing the answers.; Record the date, whether the brand was mentioned, and which websites and pages were referenced for each question.; Re-test the same set of questions monthly to track trends, rather than focusing on single results. This method is cost-effective, requiring only a simple spreadsheet and consistent execution.
First, let's discuss the limitations of this method.
Fixed query sets have clear limitations, which we explicitly state in our reports: AI responses are inherently random, and asking the same question twice may yield different results from different sources. Responses are also affected by the user's account, location, and model version; what you observe may not be representative of what all users see. Furthermore, when the model or product is updated, the entire set of benchmarks may be reset. Therefore, it can only provide information on "our brand's relative trends within this set of questions," but cannot be used to infer market share or exposure. Any report that claims "recommended by ChatGPT" based on a single screenshot should be viewed with skepticism.
What can content be used for?
Instead of chasing opaque algorithms, focus on building a verifiable foundation. This approach serves both traditional search and all AI platforms: ensuring that OAI-SearchBot can access pages (by checking robots.txt and WAF); ensuring that important content is present in plain text within HTML, rather than hidden within interactive elements; using a question-first writing structure, where a paragraph answers the question before elaborating; relying on verifiable evidence and named authors, as this increases the value of cited content; and ensuring that brand names and company information are consistent across the web, to avoid confusion.
How do you measure traffic after it has been cited?
OpenAI's official instructions state that the referral URL for ChatGPT search will automatically include "utm_source=chatgpt.com". This allows you to create a ChatGPT group in GA4 using both the campaign source and session source, and then track landing pages, case studies, CTAs, and inquiries. However, this doesn't mean that all ChatGPT traffic can be tracked. Without clicks, there will be no website data, and traffic from app openings, privacy restrictions, redirects, or parameter removals may still be categorized as referral or direct. Reports should also check both UTM and referrer data, and use inquiry numbers rather than just session counts to determine value.