This project is a Streamlit web application that leverages OpenAI's GPT-4o to generate descriptions for uploaded images
-
Updated
Jul 19, 2024 - Python
This project is a Streamlit web application that leverages OpenAI's GPT-4o to generate descriptions for uploaded images
AI Image Description Generator accurately extracts the key elements from images and interprets the creative purposes behind them, which can be applied in fields such as scientific research, artistic creation, and the mutual search between images and texts.
为 DeepSeek V4.0 等单模态大模型装上眼睛和耳朵 —— 本地视觉+音频 MCP 服务器,基于 Ollama + MiniCPM-V 4.6 + faster-whisper,图片描述·视频分析·语音转文字
Dsh-visual-plugin.Give your text-only model eyes: forward user images to any OpenAI-compatible vision model and see the results in a Web UI right panel
Image identification with Kosmos2 model, drawing and cutting bbox with object detection
Experimenting with mastodon.social client alt-text usage dataset.
DeepSeek Harness 识图插件:为不具备原生识图能力的模型提供识图能力(阿里云百炼 qwen3.5-omni-plus,失败自动切换智谱 glm-4.6v-flash)。由 claude-vision-skill 移植适配。 | Vision tool for DeepSeek Harness
Give vision to non-vision models in OpenCode using Google's free Gemini API. Transparently replaces image parts with detailed text descriptions — vision models stay untouched.
279K image alt-text pairs from 489 Bluesky accounts — curated for quality, validated at 90%+ alt-text rate
多账号 Google Gemini 提供商(DSH 插件):订阅额度即用即扣,账号池自动调度;主模型没有视觉也能贴图看图(池内 Gemini 自动转述),另有生图工具与中英双语设置页
A lightweight console utility that uses an LLM to generate descriptions and keywords for images.
给 DeepSeek Harness 纯文本模型装上原生视觉(Windows):粘贴即看图——预注入描述,模型首轮就看见,不用选模型、不用调工具;see_image 精查;自定义视觉后端(任意 OpenAI 兼容模型)+ 四后端容灾;换主模型视觉自动跟随。| Give text-only DeepSeek Harness models native-feeling vision on Windows: paste and the model just sees it — pre-injected descriptions, see_image tool, custom backends, 4-backend failover.
A new package that processes user-submitted text descriptions of images or videos containing watermarks and returns structured, watermark-free descriptions. It uses an LLM to reinterpret the content w
📊 Multi-Modal RAG 2.0 — An enterprise-grade RAG system that processes PDFs, tables, charts/images, and audio files together. Uses Unstructured.io, Camelot, Gemini Vision, Whisper, ChromaDB, FastAPI, and Streamlit. Increases information retrieval accuracy by 35% over text-only RAG.
An intelligent assistant powered by the ReAct framework, leveraging LangChain for tool-based reasoning and Gradio for a user-friendly interface. Supports tasks like weather queries, PDF summarization, image descriptions, and more.
A see_image vision tool plugin for DeepSeek Harness — describe images through any OpenAI-compatible vision model (GitHub Copilot, OpenAI, Ollama, vLLM, LM Studio).
CLI tool for generating image metadata in bulk — powered by Gemini for smarter, more context-aware descriptions.
DSH multimodal plugin: drag-image auto-describe with a configurable OpenAI-compatible vision model (host patch + agent preset + optional adapter). MIT.
It is an innovative repository housing a sophisticated Large Language Model (LLM) project, showcasing the intersection of advanced natural language processing and cutting-edge artificial intelligence. This repository serves as a comprehensive platform for the development, experimentation, and application of state-of-the-art language models.
AI-Powered-Solution-for-Assisting-Visually-Impaired-Individuals
To associate your repository with the image-description topic, visit your repo's landing page and select "manage topics."