Knowledge method
Is IT Support AI Trained on Your Internal Documentation?
One vocabulary correction with three consequences: when a fix can be used, what the model provider keeps, and why retrieval does not vouch for a runbook's accuracy.
9 min read
Updated on
A reader who searches for an “IT support AI trained on our internal documentation” is asking for training, but the implementations reviewed here, from Serval, Ravenna, Harmony and Moveworks among others, use retrieval over connected documentation: the agent searches the sources you connect when a question arrives and answers from what it finds. Nothing in these pages rules out a vendor tuning a model elsewhere. Whether a corrected article can serve depends first on two conditions. The sync must have succeeded: Serval syncs every four hours, Ravenna every 24 hours by default, and Harmony shows a failed status per article when it does not. The article must be in a usable state, published or visible rather than draft, hidden or archived. Verify separately whether the connector applies the requester’s permissions in the source: some documentation describes a scope per folder, channel or agent instead. Even when all of this holds, a revision becomes available for retrieval; nothing guarantees that it is used in the next answer.
Does an AI helpdesk that reads our Notion docs learn them, or search them?
The products read for this guide search them. Ravenna’s knowledge documentation draws the line in one sentence: “Without knowledge, the agent relies solely on its base training and rules.” With knowledge connected, the agent searches the folders it has been given and ranks the results by relevance. Serval syncs connected content read-only, and its help desk agent “searches and cites it”. Moveworks splits documents into chunks, turns each into a vector embedding and retrieves from that index.
This describes retrieval, not a training run on your pages. Two vendors use the verb train for it: Harmony’s SharePoint page offers to use SharePoint content “to train Harmony AI”, content Harmony “fetches and indexes”, and Atomicwork’s Notion page says you can “train Atom on your Notion documents”. Whether customer content trains a model anywhere is a separate question, taken up below. Retrieval also changes what a missing answer looks like. Siit’s trust model states “You choose the knowledge sources” and “Answers cite sources”, and where no trustworthy source exists the agent “defaults to ‘I don’t know’” and can open or route a request.
An agent that reads your documentation can still answer from elsewhere. Harmony’s default, “Knowledge Base + trained data mode”, supplements your content with general trained knowledge when needed; its knowledge-base-only mode escalates to a human “rather than generating a response from general training data”.
Serval and Ravenna both list Notion among their sources. What each source exposes is covered in what an agent reads once a knowledge base is connected. Drive-based documentation has its own page: Google Drive as a knowledge source.
When can a corrected article be used in an answer?
No retraining step stands between an edit and the agent, but two gates do, the sync and the state of the document, and a third question applies: whether the connector enforces the asker’s permissions in the source.
The sync.
| Product | Published sync behavior | Other way in |
|---|---|---|
| Serval | “connected sources automatically sync every four hours” | Immediate manual sync |
| Ravenna | “Knowledge folders auto-sync every 24 hours by default” | “Manual sync is available at any time” |
| Atomicwork (Google Drive) | Sync every 24 hours | “Sync now” for immediate updates |
| Moveworks (indexed connectors) | Full sync daily, incremental every 15 minutes by default; its ingestion schedule page gives every four hours for a knowledge article, 48 hours at worst | Webhook updates where the source supports them |
| Zendesk | External content as of the last sync, “which usually happens every 24 hours” | Help center articles are searched at the time of the request |
| Harmony | “Articles sync periodically”, no interval published | Status per article (Completed, Pending, Failed), re-sync after a failure |
A sync can also fail: Ravenna says sync “requires active integration authentication”, and Harmony cites permission changes in the source, deleted content and integration issues.
The state. Serval gives each synced document a Visible or Hidden status, and a hidden document “is synced but not used by the agent”; pages authored in Serval follow a draft and publish model, and “unpublished pages are never searched by the agent.” Ravenna excludes archived documents from agent searches. Freshworks says its AI Agent “learns directly from your published articles”. On deleted documents, Ravenna’s pages differ: its knowledge overview says it “marks it with an error but does not delete it”, its Notion sync page that “Deleted pages are removed from your knowledge base.”
The permissions. Serval says it respects source permissions “when the integration supports propagating access controls”, and that “The agent only surfaces a document to a user who is allowed to see it in the source system.” Zendesk says AI agent responses respect article view permissions, and that an unauthenticated customer gets answers from public articles only. Where a connector enforces source permissions, a correction in a restricted page reaches only the people who could read it; where it scopes by folder, channel or agent instead, check what that scope admits.
Put together: after a successful sync, an eligible and visible revision can become available for retrieval; this does not guarantee its use in the next answer, since retrieval ranks candidate passages by relevance. The delay on your stack is something to measure.
What does a no-training commitment cover, and what is kept?
Retrieved documentation is sent to a language model as input, so the model provider’s terms belong in the evaluation. A request for ChatGPT for internal IT support ends up here too: an agent built on OpenAI’s API falls under the API terms read below, a page that does not describe the terms of ChatGPT’s own plans. Siit’s trust model says it uses model providers “through ‘no training’ endpoints when available” and does not permit providers “to use your content for their own model training”. It names OpenAI and Mistral AI; it says nothing about retention, and “when available” leaves open which calls qualify. The providers publish what such a commitment contains, and it is narrower than a promise that nothing is kept.
OpenAI states: “As of March 1, 2023, data sent to the OpenAI API is not used to train or improve OpenAI models (unless you explicitly opt in to share data with us).” Abuse monitoring logs, which may contain prompts and responses, are “retained for up to 30 days” by default, longer if law or harm prevention requires it. Zero Data Retention needs prior approval by OpenAI, and “Zero Data Retention excludes customer content from abuse monitoring logs in the same way as Modified Abuse Monitoring.” Eligibility is decided per endpoint: in October 2026, /v1/chat/completions, /v1/responses and /v1/embeddings are eligible, the first two with limitations, while /v1/conversations, /v1/assistants and /v1/vector_stores are not, and “may retain application state when used, even if you have Zero Data Retention enabled.”
Anthropic writes: “Retained data is never used for model training without your express permission.” Its Zero Data Retention arrangement is enabled per organization on request, and under it “Anthropic does not store customer prompts or responses at rest after the API response is returned.” Listed exceptions include content flagged by its automated trust and safety systems, which may be kept for up to 2 years.
Mistral AI states that “ZDR and training opt-out are separate controls”. Zero data retention applies to supported stateless API calls on paid plans, including chat completions and embeddings, and not to Agents, Conversations, Libraries or the Files API: “ZDR does not apply to stateful APIs or products.”
These are three statements, not one. On training, OpenAI states a default for API data, Anthropic’s sentence covers retained data, and Mistral AI states no default on that page. Retention is a separate regime that depends on the endpoint, the approval and the arrangement in force, and none of these pages says which arrangement your vendor holds. That answer belongs in the audit file: what a compliance review of an IT agent examines.
The answers are worth what the documentation is worth
Retrieval does not establish that a source is accurate or current: an agent may present a runbook for a VPN client retired last year as confidently as an article reviewed this morning.
That moves the work from training to editing. Scope narrowly first: Ravenna advises focused folders rather than one large collection, and Serval says to sync only what the agent should reference: “Exclude drafts, internal notes, or outdated content.” Resolve duplicates and contradictions in the source. Treat publication as a review. Investigate missing or incorrect answers across source content, synchronization, permissions, retrieval and agent configuration before deciding what to change. A declined answer is not a ticket the agent prevented: what a deflection number covers.
Retrieval, not a training run
Ravenna's documentation draws the line: without knowledge, the agent relies on its base training. With knowledge, it searches the folders you connected and ranks the results by relevance.
An edit waits for a successful sync
Serval documents an automatic sync every four hours and Ravenna every 24 hours by default, each with a manual sync. Harmony shows a Completed, Pending or Failed status per article.
No training is not no retention
OpenAI states that API data has not been used for training since March 2023 unless a customer opts in, while abuse monitoring logs are kept for up to 30 days by default, longer if law or harm prevention requires it.
A refusal is a starting point
Where no source answers, Siit's agent says it does not know and can route a request, and Harmony's knowledge-base-only mode escalates to a human. Each case is worth investigating: the source content, the sync, permissions, retrieval or the agent's configuration may be the cause.
A first evaluation cycle on your own runbooks
The following is a proposed first evaluation cycle; its duration depends on access approvals and on the state of the corpus.
- Start from the ten requests the desk answers most often, taken from the queue.
- Connect one collection, the one you would defend in a review.
- Test the update path on one question: edit the article, trigger the manual sync where offered, check that the sync succeeded and that the revision is published or visible, then ask again, and once more later.
- Ask the ten questions from two test accounts with different permissions, and read the cited source behind each answer.
- Apply that investigation to each missing or incorrect answer before changing anything, then widen the scope.
- Ask the vendor in writing which model providers handle your content, on which endpoints, under which retention terms, and whether customer content is used to train or tune any model.
The rollout around an existing ITSM is covered in adding an agent without replacing your ITSM.
Which vendors document where their agent’s answers come from?
Listed here: vendors of an AI service desk agent whose public documentation, read for this guide in October 2026, states which knowledge sources the agent answers from and how they sync (a frequency or a trigger), and which content it leaves out by state or permissions.
- Atlassian: Its Teamwork Graph connector page says content created or removed in a connected app such as Google Drive, Confluence or SharePoint is synced with Rovo, and restricted data stays visible only to users with access in that app.
- Freshworks: Its AI Agent knowledge sources page lists URLs, files, solution articles and apps, says URLs and solution articles “re-sync automatically”, and that the agent “honors the visibility settings” of solution articles.
- Harmony: Its knowledge base settings page lists Confluence, Notion, Freshservice and SharePoint, shows a Completed, Pending or Failed status per article with a re-sync after a failure, and says “Permissions come from the source system.”
- Ravenna: Its knowledge overview lists sources including Notion, Google Drive, Confluence and Slack, describes a 24-hour default sync with manual sync, and says “Archived documents are excluded from agent searches.”
- Risotto: Its product pages say it answers from Notion, Confluence, Drive and Slack, keeps answers current “automatically as your docs change”, and “respects Notion permissions and access rules.”
- Serval: Its connected sources page lists Notion, Confluence and Google Drive among others, a four-hour sync with immediate sync, a Hidden status the agent ignores, and source permissions where the integration propagates them.
- Siit: Its Confluence page describes a scheduled sync with a manual Synchronize action and drafts excluded from AI; its AI trust model says the AI “only sees data the requester is allowed to see.”
- Zendesk: Its page on AI agent knowledge sources says help center content is searched live, external content as of the last sync, usually every 24 hours, under article view permissions.
Frequently asked questions
Can we get an IT support AI trained on our internal documentation?
The implementations reviewed in this guide use retrieval over connected documentation, which is different from training a model on that documentation. The agent searches the sources you connect when a question arrives and answers from what it finds, as Serval, Ravenna and Moveworks document. That does not rule out a vendor tuning a model elsewhere: ask. A correction then depends on a successful sync, a published or visible state and, depending on the connector, the asker's permissions in the source, not on a retraining cycle.
How does an AI helpdesk that reads our Notion docs stay current?
Through a sync of the connected workspace. Serval documents an automatic sync every four hours and Ravenna every 24 hours by default, both with a manual sync. After a successful sync, an eligible and visible revision can become available for retrieval; nothing guarantees it is used in the next answer.
Can we use ChatGPT for internal IT support on our own documentation?
This guide read OpenAI's API data page, which applies to agents built on OpenAI's API: OpenAI states that data sent to its API is not used for training unless a customer opts in, and that abuse monitoring logs are kept for up to 30 days by default; Zero Data Retention needs OpenAI's approval and covers listed endpoints only. That API page does not describe the terms of ChatGPT's own plans: if staff use ChatGPT directly, read those separately. Ask the vendor which provider, endpoints and retention terms apply to your tenant.
Which vendors document where their agent's answers come from and how they stay current?
Among the vendors read for this guide in October 2026, those whose public documentation states which knowledge sources the agent answers from and how they sync (a frequency or a trigger), and which content it leaves out by state or permissions: Atlassian, Freshworks, Harmony, Ravenna, Risotto, Serval, Siit and Zendesk. The list ranks none of them.
Sources
- AI trust model: knowledge sources, citations, no-training endpoints and model providers · Siit
- Confluence: scheduled sync, Synchronize action and drafts excluded from AI · Siit
- Your data: training, abuse monitoring retention and Zero Data Retention eligibility · OpenAI
- API and data retention: Zero Data Retention scope and retention commitments · Anthropic
- Zero data retention: covered endpoints and difference from training opt-out · Mistral AI
- Connect to external sources: supported sources, four-hour sync, visibility and permissions · Serval
- Knowledge overview: hybrid search, 24-hour sync and archived documents · Ravenna
- Notion knowledge sync: auto-sync of edits, deleted pages removed · Ravenna
- Managing the knowledge base: sync status, source permissions and general AI knowledge setting · Harmony
- SharePoint integration: content fetched and indexed to train Harmony AI · Harmony
- How Moveworks ingests content: sync types, chunking, embeddings and permissions · Moveworks
- Moveworks Ingestion Schedule: knowledge articles every four hours, 48 hours worst case · Moveworks
- Connecting knowledge sources to power generative replies in AI agents · Zendesk
- Connect your knowledge sources to the AI Agent · Freshworks
- Google Drive: 24-hour sync, Sync now and inherited permissions · Atomicwork
- Notion integration: train Atom on your Notion documents · Atomicwork
- How Teamwork Graph connector permissions are kept in sync · Atlassian
- Risotto product page: answers from Notion, Confluence, Drive and Slack, kept current as docs change · Risotto
- Risotto product page: Notion integration that respects Notion permissions · Risotto