Multimodal★ 3.5k
Convert pasted images into structured OCR, layout, entity, and relationship evidence so text-only DeepSeek or GLM models can work with screenshots and document images.
Interfaces, tools, and workflows that extend DeepSeek Harness
Convert pasted images into structured OCR, layout, entity, and relationship evidence so text-only DeepSeek or GLM models can work with screenshots and document images.
Provide agent-callable tools for image Q&A, long-screenshot OCR, crop, pixel and color checks, UI reconstruction, and visual diffs.
Open the WeShop infinite canvas inside DeepSeek Harness so agents can generate and organize product visuals while synchronizing work between the session and canvas.