← back to blog
    Tom StrömSeptember 29, 20265 min read

    Attach Files and Screenshots in Cogny Chat — On Every Plan

    You could already paste a screenshot into the Cogny chat — since 23 September, when a customer tried to hand us a Looker Studio dashboard to rebuild and found there was no way to. What that release didn't say out loud is that the screenshot only reached the model on workspaces running Claude. Every Free, Solo, Cogny AI, Cogny Pro and Cloud workspace defaults to DeepSeek V4 Flash for chat (that's what the credit allowances are priced against), and DeepSeek takes no image input. On that path the upload succeeded, the thumbnail rendered, and the agent answered as if you'd sent nothing.

    That is fixed, and while we were in there we added documents. From today the chat composer takes:

    • Images: any image your browser can open. PNG, JPEG, GIF and WebP go up as they are; HEIC from an iPhone, SVG, BMP, TIFF and AVIF are converted to PNG in the browser before upload, and anything over 10 MB is downscaled rather than refused. You'll see a "Converted" note when that happened.
    • Documents: PDF, DOCX, CSV, TSV, TXT, Markdown, JSON, HTML, up to 25 MB.

    Up to five per message, on every plan, by paste, drag-and-drop or the paperclip. You can send an attachment with no caption at all; the thread gets titled after the filename.

    What the agent actually sees

    The interesting part is not the upload, it's what the model gets. Three cases.

    Documents, on every provider. No model ever receives the file bytes. When you attach a PDF, the text is extracted at upload time (pdf-parse for PDFs, mammoth for DOCX, a passthrough for the text formats) and stored next to the original. When the agent runs, the text is inlined into your message inside an <attached_file name="…" type="…" size="…"> element. The same text goes to Claude, GPT-5.5 and DeepSeek, so a brief reads identically no matter which model your workspace runs. Per file the agent gets up to 80,000 characters (about 20,000 tokens); past that it's told exactly how many characters were cut, so it says "I read the first 30 pages" instead of pretending it read all 90.

    Images, on a model that can see. Claude and OpenAI workspaces get the image as an image block, as before.

    Images, on a model that can't. DeepSeek and Berget workspaces now get a transcription instead of nothing. Once per image, a vision model (Claude Opus 5.5, low effort) writes a literal transcription: every visible label and number, tables as tables, chart series with the values it can read off them, what kind of thing the image is and how it's laid out. It's asked not to summarize or advise, just to transcribe, so your chat model does the reasoning with the same facts a sighted model would have. The transcription is cached on the message, so replaying the thread on later turns costs nothing extra. The agent's own view_image tool takes the same route on those providers, so "look at this URL" works too.

    Attachments sent while the agent is mid-run are picked up between steps the same way: documents inlined, images transcribed or handed to view_image.

    It stays in your context tree

    An attachment is not a one-thread thing. Every file and image you attach is saved to the workspace's context tree the moment it's uploaded, so the brief you handed the agent on Monday is there for the scheduled report on Friday and for every teammate's chat in between.

    • Documents land under uploads/chat/<name> with their full extracted text as the node's content, the same shape as a document uploaded from Settings → Context Tree. The agent is told the path in the message, so if the inline excerpt was truncated it can read the rest with read_context_node.
    • Images land under assets/uploads/<name>, the same place the image-gen tools keep uploads, so list_uploads finds them and generate_image can use them as references. When a text-only workspace triggers a transcription, that text is appended to the node too, which makes the screenshot searchable.
    • The conversations/<thread> node in the tree now lists what was attached in each message and links the node, and a thread that started with an attachment is titled after the filename instead of "Untitled conversation".

    If something was a one-off you don't want kept, archive the node from Settings → Context Tree; the chat keeps working from the message.

    What it costs

    • Document extraction is free. It's CPU on our side, no inference.
    • An image on a Claude or OpenAI workspace is billed as part of the turn's input tokens, as before.
    • An image on a DeepSeek or Berget workspace costs one vision-model call, a few cents, logged to your credit ledger under chat_attachment_image_description. That's once per image, not once per turn.

    Where the original files live

    Images go to a public bucket with an unguessable path, because the model API has to be able to fetch them on every replay of the thread. Documents are different: briefs, contracts and exports carry customer data, and no API needs them by URL. They live in a private bucket that only Cogny's server reads after checking you're a member of the workspace, and the chip in your transcript opens the original through a five-minute signed link.

    What's not there yet

    • Spreadsheets. .xlsx isn't parsed; export the sheet as CSV. The composer says so when you try.
    • Scanned PDFs. If a PDF has no text layer, the upload is refused with the reason and you're pointed at screenshotting the pages, which the agent reads visually.
    • Images your browser can't decode. Conversion happens client-side, so a HEIC in Chrome on Windows (which can't open HEIC) is refused by name; Safari and iOS handle it. Animated GIFs over 10 MB are refused too, since re-encoding would drop the animation.
    • Video and audio. Not in chat. The Ad Studio has its own video upload.

    Source of truth for the accepted types is src/lib/chat/attachments.ts; the provider routing lives in src/lib/chat/attachment-context.ts.