Skip to content

Commit 0fc4524

Browse files
committed
feat(confluence): index PowerPoint and Excel attachments
Extend the Confluence attachment allowlist so .pptx and .xlsx files on synced pages and blog posts are listed and handed to the shared parser pipeline the same way PDF and Word attachments already are. Macro-enabled, template, legacy binary and OpenDocument variants stay excluded. Add listing, hydration, genuine-bytes roundtrip and renamed-to-unsupported coverage, and update the connector guides to name the new formats.
1 parent b0a68f2 commit 0fc4524

4 files changed

Lines changed: 119 additions & 38 deletions

File tree

‎apps/docs/content/docs/knowledgebase/connectors.mdx‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -78,7 +78,7 @@ Each connector has source-specific fields that control what gets synced. Example
7878

7979
- **Notion** — sync an entire workspace, a specific database, or a single page tree
8080
- **GitHub** — specify a repository, branch, and optional file extension filter
81-
- **Confluence** — enter your Atlassian domain and choose spaces, or **All** for all spaces accessible at each sync. Optionally filter by content type or label. PDF and Word (`.docx`, Word 97–2003 `.doc`) attachments on matching pages and blog posts are included as separate documents.
81+
- **Confluence** — enter your Atlassian domain and choose spaces, or **All** for all spaces accessible at each sync. Optionally filter by content type or label. PDF, Word (`.docx`, Word 97–2003 `.doc`), Excel (`.xlsx`), and PowerPoint (`.pptx`) attachments on matching pages and blog posts are included as separate documents.
8282
- **Azure DevOps** — choose what to sync (wiki pages, work items, repository files, or all), with optional work item type/state filters, a custom WIQL query, and repository/branch/path filters
8383
- **Amazon S3** — point at a bucket with an optional key prefix and a customizable file extension allowlist; S3-compatible stores (Cloudflare R2, MinIO) are supported via a custom endpoint
8484
- **YouTube** — sync a channel (by `@handle` or ID) or playlist, with an optional published-after date filter and the option to exclude Shorts
@@ -88,7 +88,7 @@ Each connector has source-specific fields that control what gets synced. Example
8888

8989
Configuration is validated on save — if a repository doesn't exist or a domain is unreachable, you'll see an error immediately.
9090

91-
Confluence attachment indexing requires `read:attachment:confluence`. For a service account, include it when creating the scoped API token; see the [Confluence scope list](/search/confluence#using-a-service-account). Attachments are checked even when the parent page has not changed. Files over 100 MB appear as skipped; convert Word 6/95 files to `.docx` before attaching them.
91+
Confluence attachment indexing requires `read:attachment:confluence`. For a service account, include it when creating the scoped API token; see the [Confluence scope list](/search/confluence#using-a-service-account). Attachments are checked even when the parent page has not changed. Files over 100 MB appear as skipped; convert Word 6/95 files to `.docx`, and `.xls` and `.ppt` files to `.xlsx` and `.pptx`, before attaching them. Spaces that were already connected pick up newly supported formats on their next sync.
9292

9393
</Step>
9494
<Step>

‎apps/docs/content/docs/search/confluence.mdx‎

Lines changed: 4 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -7,7 +7,7 @@ import { Callout } from 'fumadocs-ui/components/callout'
77
import { Step, Steps } from 'fumadocs-ui/components/steps'
88
import { Image } from '@/components/ui/image'
99

10-
Search pages, blog posts, and their PDF and Word attachments from selected Confluence Cloud spaces. A Sim organization admin enables Confluence; **each teammate connects their own account**.
10+
Search pages, blog posts, and their PDF, Word, Excel, and PowerPoint attachments from selected Confluence Cloud spaces. A Sim organization admin enables Confluence; **each teammate connects their own account**.
1111

1212
| Method | How it works |
1313
| --- | --- |
@@ -125,9 +125,9 @@ See Atlassian's [account setup](https://support.atlassian.com/user-management/do
125125
| **Filter by Label** | Optional comma-separated labels; content can match any listed label. |
126126
| **Metadata tags** | Labels, version, and last-modified tags. |
127127

128-
Search manages the schedule and hides item limits. It indexes published/current content and each page's own text, including supported local callouts and code blocks. PDF, Word `.docx`, and Word 97–2003 `.doc` attachments on the selected pages and blog posts are indexed as separate documents with their parent content's permissions. Space, content-type, and label filters apply to the parent content. Attachment changes are checked on each sync, even when the parent text has not changed.
128+
Search manages the schedule and hides item limits. It indexes published/current content and each page's own text, including supported local callouts and code blocks. PDF, Word `.docx`, Word 97–2003 `.doc`, Excel `.xlsx`, and PowerPoint `.pptx` attachments on the selected pages and blog posts are indexed as separate documents with their parent content's permissions. Space, content-type, and label filters apply to the parent content. Attachment changes are checked on each sync, even when the parent text has not changed.
129129

130-
Archived content, comments, other attachment formats, and expanded Include Page, Excerpt Include, or third-party macro output are excluded. Referenced pages can be indexed separately with their own permissions. Attachments over 100 MB are shown as skipped; convert older Word 6/95 files to `.docx` before attaching them.
130+
Archived content, comments, other attachment formats, and expanded Include Page, Excerpt Include, or third-party macro output are excluded. Referenced pages can be indexed separately with their own permissions. Attachments over 100 MB are shown as skipped; convert older Word 6/95 files to `.docx`, and `.xls` and `.ppt` files to `.xlsx` and `.pptx`, before attaching them. Spaces that were already connected pick up newly supported formats on their next sync.
131131

132132
## Manage access and sync
133133

@@ -151,7 +151,7 @@ In **Sync history**, **Continuing** means a healthy listing needs another batch.
151151
| A new page, blog post, or label is missing | Confluence search can take time to update. Once the content appears in Confluence search with the selected label, sync again. |
152152
| A restricted page is missing | Both your account and the crawling account need access to the page and its ancestors. |
153153
| Embedded content is missing | Index the referenced page separately; remote macro output is excluded. |
154-
| PDF or Word attachments are missing | Check `read:attachment:confluence` and access to the parent page. Existing service-account tokens may need to be replaced with one that includes this scope. Attachment access failures are reported as a partial sync. |
154+
| Attachments are missing | Check `read:attachment:confluence` and access to the parent page. Existing service-account tokens may need to be replaced with one that includes this scope. Attachment access failures are reported as a partial sync. |
155155
| **Reconnect** or email mismatch | Authorize with the Atlassian account matching your verified Sim email and grant all requested permissions. |
156156

157157
Open a missing page as the affected teammate, check its space and page restrictions, then sync again after correcting access. See Atlassian's [content access](https://support.atlassian.com/confluence-cloud/docs/add-or-remove-page-restrictions/) and [permission inspection](https://support.atlassian.com/confluence-cloud/docs/inspect-a-users-permissions/) guides.

‎apps/sim/connectors/confluence/attachments.test.ts‎

Lines changed: 104 additions & 29 deletions
Original file line numberDiff line numberDiff line change
@@ -4,12 +4,13 @@
44
import JSZip from 'jszip'
55
import { PDFDocument, StandardFonts } from 'pdf-lib'
66
import { afterEach, beforeEach, describe, expect, it, vi } from 'vitest'
7+
import * as XLSX from 'xlsx'
78
import { DEFAULT_MAX_ERROR_BODY_BYTES, PayloadSizeLimitError } from '@/lib/core/utils/stream-limits'
89
import { parseBuffer } from '@/lib/file-parsers'
910
import { listConfluenceAttachments } from '@/connectors/confluence/attachments'
1011
import { confluenceConnector } from '@/connectors/confluence/confluence'
1112
import type { ExternalDocument, ExternalDocumentList } from '@/connectors/types'
12-
import { CONNECTOR_MAX_FILE_BYTES } from '@/connectors/utils'
13+
import { CONNECTOR_MAX_FILE_BYTES, PIPELINE_PARSED_MIME_TYPES } from '@/connectors/utils'
1314

1415
const { secureDownload } = vi.hoisted(() => ({ secureDownload: vi.fn() }))
1516
vi.mock('@/lib/knowledge/documents/secure-fetch.server', async (importOriginal) => ({
@@ -90,6 +91,75 @@ async function get(externalId = 'attachment:page:p1:att123', config = CONFIG) {
9091
return confluenceConnector.getDocument('token', config, externalId, { ...CONTEXT })
9192
}
9293

94+
const ROUNDTRIP_TEXT = 'Confluence attachment text'
95+
const OOXML_REL = 'http://schemas.openxmlformats.org/officeDocument/2006/relationships'
96+
const PACKAGE_RELS_NS = 'xmlns="http://schemas.openxmlformats.org/package/2006/relationships"'
97+
const PRESENTATION_NS =
98+
'xmlns:a="http://schemas.openxmlformats.org/drawingml/2006/main" xmlns:p="http://schemas.openxmlformats.org/presentationml/2006/main" xmlns:r="http://schemas.openxmlformats.org/officeDocument/2006/relationships"'
99+
100+
async function pdfBytes(): Promise<Buffer> {
101+
const document = await PDFDocument.create()
102+
const font = await document.embedFont(StandardFonts.Helvetica)
103+
document.addPage().drawText(ROUNDTRIP_TEXT, { font, size: 14 })
104+
return Buffer.from(await document.save())
105+
}
106+
107+
async function docxBytes(): Promise<Buffer> {
108+
const zip = new JSZip()
109+
zip.file(
110+
'[Content_Types].xml',
111+
'<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types"><Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/><Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/></Types>'
112+
)
113+
zip.file(
114+
'_rels/.rels',
115+
`<Relationships ${PACKAGE_RELS_NS}><Relationship Id="rId1" Type="${OOXML_REL}/officeDocument" Target="word/document.xml"/></Relationships>`
116+
)
117+
zip.file(
118+
'word/document.xml',
119+
`<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"><w:body><w:p><w:r><w:t>${ROUNDTRIP_TEXT}</w:t></w:r></w:p></w:body></w:document>`
120+
)
121+
return zip.generateAsync({ type: 'nodebuffer' })
122+
}
123+
124+
/** One slide resolved through `p:sldIdLst`, the order the presentation walker follows. */
125+
async function pptxBytes(): Promise<Buffer> {
126+
const zip = new JSZip()
127+
zip.file(
128+
'[Content_Types].xml',
129+
'<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types"/>'
130+
)
131+
zip.file(
132+
'ppt/presentation.xml',
133+
`<?xml version="1.0" encoding="UTF-8" standalone="yes"?><p:presentation ${PRESENTATION_NS}><p:sldIdLst><p:sldId id="256" r:id="rId1"/></p:sldIdLst></p:presentation>`
134+
)
135+
zip.file(
136+
'ppt/_rels/presentation.xml.rels',
137+
`<Relationships ${PACKAGE_RELS_NS}><Relationship Id="rId1" Type="${OOXML_REL}/slide" Target="slides/slide1.xml"/></Relationships>`
138+
)
139+
zip.file(
140+
'ppt/slides/slide1.xml',
141+
`<?xml version="1.0" encoding="UTF-8" standalone="yes"?><p:sld ${PRESENTATION_NS}><p:cSld><p:spTree><p:sp><p:nvSpPr><p:cNvPr id="1" name="s"/><p:cNvSpPr/></p:nvSpPr><p:txBody><a:bodyPr/><a:p><a:r><a:t>${ROUNDTRIP_TEXT}</a:t></a:r></a:p></p:txBody></p:sp></p:spTree></p:cSld></p:sld>`
142+
)
143+
return zip.generateAsync({ type: 'nodebuffer' })
144+
}
145+
146+
async function xlsxBytes(): Promise<Buffer> {
147+
const book = XLSX.utils.book_new()
148+
XLSX.utils.book_append_sheet(
149+
book,
150+
XLSX.utils.aoa_to_sheet([['Note'], [ROUNDTRIP_TEXT]]),
151+
'Sheet1'
152+
)
153+
return XLSX.write(book, { type: 'buffer', bookType: 'xlsx' }) as Buffer
154+
}
155+
156+
const ROUNDTRIP_FIXTURES = {
157+
pdf: pdfBytes,
158+
docx: docxBytes,
159+
pptx: pptxBytes,
160+
xlsx: xlsxBytes,
161+
} as const
162+
93163
beforeEach(() => {
94164
fetchMock.mockReset()
95165
secureDownload.mockReset().mockResolvedValue(new Response('binary bytes'))
@@ -98,15 +168,20 @@ beforeEach(() => {
98168
afterEach(() => vi.unstubAllGlobals())
99169

100170
describe('Confluence attachment listing', () => {
101-
it('lists PDF, DOC and DOCX stubs without downloading, and excludes unsupported or archived files', async () => {
171+
it('lists PDF, Word, PowerPoint and Excel stubs without downloading, and excludes unsupported or archived files', async () => {
102172
fetchMock.mockResolvedValue(
103173
Response.json({
104174
results: [
105175
file(),
106176
file({ id: '2', title: 'Legacy.DOC' }),
107177
file({ id: '3', title: 'Modern.docx' }),
108-
file({ id: '4', title: 'image.png' }),
109-
file({ id: '5', title: 'old.pdf', status: 'archived' }),
178+
file({ id: '4', title: 'Deck.pptx' }),
179+
file({ id: '5', title: 'Sheet.XLSX' }),
180+
file({ id: '6', title: 'image.png' }),
181+
file({ id: '7', title: 'Macro.pptm' }),
182+
file({ id: '8', title: 'Legacy.xls' }),
183+
file({ id: '9', title: 'Old.ppt' }),
184+
file({ id: '10', title: 'old.pdf', status: 'archived' }),
110185
],
111186
})
112187
)
@@ -119,7 +194,14 @@ describe('Confluence attachment listing', () => {
119194
'attachment:page:p1:att123',
120195
'attachment:page:p1:2',
121196
'attachment:page:p1:3',
197+
'attachment:page:p1:4',
198+
'attachment:page:p1:5',
122199
])
200+
expect(result.documents.slice(1).map((doc) => doc.mimeType)).toEqual(
201+
['pdf', 'doc', 'docx', 'pptx', 'xlsx'].map((extension) =>
202+
PIPELINE_PARSED_MIME_TYPES.get(extension)
203+
)
204+
)
123205
expect(
124206
result.documents.slice(1).every((doc) => doc.contentDeferred && doc.content === '')
125207
).toBe(true)
@@ -364,7 +446,7 @@ describe('Confluence attachment hydration', () => {
364446
}
365447
)
366448

367-
it.each(['pdf', 'doc', 'docx'])(
449+
it.each(['pdf', 'doc', 'docx', 'pptx', 'xlsx'])(
368450
'hands an original %s file to the shared parser pipeline',
369451
async (extension) => {
370452
fixture(file({ title: `Guide.${extension}` }))
@@ -492,37 +574,30 @@ describe('Confluence attachment hydration', () => {
492574
expect(secureDownload.mock.calls[0][1].signal).toBe(signal)
493575
})
494576

495-
it.each(['pdf', 'docx'] as const)(
577+
it.each(['pdf', 'docx', 'pptx', 'xlsx'] as const)(
496578
'roundtrips genuine %s bytes through the public parser',
497579
async (extension) => {
498-
let bytes: Buffer
499-
if (extension === 'pdf') {
500-
const document = await PDFDocument.create()
501-
const font = await document.embedFont(StandardFonts.Helvetica)
502-
document.addPage().drawText('Confluence attachment text', { font, size: 14 })
503-
bytes = Buffer.from(await document.save())
504-
} else {
505-
const zip = new JSZip()
506-
zip.file(
507-
'[Content_Types].xml',
508-
'<Types xmlns="http://schemas.openxmlformats.org/package/2006/content-types"><Default Extension="rels" ContentType="application/vnd.openxmlformats-package.relationships+xml"/><Override PartName="/word/document.xml" ContentType="application/vnd.openxmlformats-officedocument.wordprocessingml.document.main+xml"/></Types>'
509-
)
510-
zip.file(
511-
'_rels/.rels',
512-
'<Relationships xmlns="http://schemas.openxmlformats.org/package/2006/relationships"><Relationship Id="rId1" Type="http://schemas.openxmlformats.org/officeDocument/2006/relationships/officeDocument" Target="word/document.xml"/></Relationships>'
513-
)
514-
zip.file(
515-
'word/document.xml',
516-
'<w:document xmlns:w="http://schemas.openxmlformats.org/wordprocessingml/2006/main"><w:body><w:p><w:r><w:t>Confluence attachment text</w:t></w:r></w:p></w:body></w:document>'
517-
)
518-
bytes = await zip.generateAsync({ type: 'nodebuffer' })
519-
}
580+
const bytes = await ROUNDTRIP_FIXTURES[extension]()
520581
fixture(file({ title: `Guide.${extension}`, fileSize: bytes.length }))
521582
secureDownload.mockResolvedValue(new Response(bytes))
522583
const doc = await get()
523584
expect(doc?.sourceFile).toBeDefined()
524585
const parsed = await parseBuffer(doc!.sourceFile!.bytes, extension)
525-
expect(parsed.content).toContain('Confluence attachment text')
586+
expect(parsed.content).toContain(ROUNDTRIP_TEXT)
587+
}
588+
)
589+
590+
it.each(['Guide.ppt', 'Guide.xls'])(
591+
'replaces an attachment renamed to unsupported %s without downloading',
592+
async (title) => {
593+
fixture(file({ title }))
594+
const doc = await get()
595+
expect(doc?.skippedReason).toBe(
596+
'Attachment is no longer a PDF, Word, Excel or PowerPoint document'
597+
)
598+
expect(doc?.skippedExistingDisposition).toBe('replace')
599+
expect(doc?.contentDeferred).toBe(false)
600+
expect(secureDownload).not.toHaveBeenCalled()
526601
}
527602
)
528603
})

‎apps/sim/connectors/confluence/attachments.ts‎

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -26,7 +26,13 @@ const PAGE_SIZE = 50
2626
const REQUESTS_PER_CALL = 5
2727
const MAX_CURSOR_BYTES = 512 * 1024
2828
const MAX_METADATA_BYTES = 2 * 1024 * 1024
29-
const FILE_EXTENSIONS = new Set(['pdf', 'doc', 'docx'])
29+
/**
30+
* Attachment formats listed for indexing: a deliberate subset of the shared
31+
* `PIPELINE_PARSED_MIME_TYPES`, limited to the headline PDF, Word, Excel and
32+
* PowerPoint extensions. Macro-enabled, template, legacy binary (`.xls`, `.ppt`)
33+
* and OpenDocument variants are not listed for Confluence.
34+
*/
35+
const FILE_EXTENSIONS = new Set(['pdf', 'doc', 'docx', 'pptx', 'xlsx'])
3036
const boundedId = z.string().min(1).max(254)
3137
const providerCursor = z.string().min(1).max(8192)
3238
const parentSchema = z.object({ id: boundedId, type: z.enum(['page', 'blogpost']) })
@@ -406,7 +412,7 @@ function assertDownloadUrl(value: string): void {
406412
}
407413
}
408414

409-
/** Downloads a version-pinned original for the shared PDF/OCR and Word parsing pipeline. */
415+
/** Downloads a version-pinned original for the shared PDF/OCR, Word, Excel and PowerPoint parsing pipeline. */
410416
export async function getConfluenceAttachment(
411417
input: AttachmentRequest,
412418
sourceConfig: Record<string, unknown>,
@@ -420,7 +426,7 @@ export async function getConfluenceAttachment(
420426
const mimeType = attachmentMimeType(attachment)
421427
if (!mimeType) {
422428
return {
423-
...markSkipped(stub, 'Attachment is no longer a PDF or Word document'),
429+
...markSkipped(stub, 'Attachment is no longer a PDF, Word, Excel or PowerPoint document'),
424430
skippedExistingDisposition: 'replace',
425431
}
426432
}

0 commit comments

Comments
 (0)