Skip to content
#

vision-models

Here are 67 public repositories matching this topic...

A curated collection of resources focused on the Mechanistic Interpretability (MI) of Large Multimodal Models (LMMs). This repository aggregates surveys, blog posts, and research papers that explore how LMMs represent, transform, and align multimodal information internally.

  • Updated Mar 4, 2026
computer-vision-challenge

Generative AI document analysis on Amazon Bedrock AgentCore: a Strands agent routes pages through 27 vision specialists via an MCP Gateway, then correlates results into a structured, accessibility-oriented document tree. Eight-model registry (Anthropic, Amazon, OpenAI), specialist wizard, scripted CDK deployment, React UI on ECS with Cognito.

  • Updated Sep 18, 2026
  • Python

Enhance your skills in prompt engineering for vision models. Learn to effectively prompt, fine-tune, and track experiments for models like SAM, OWL-ViT, and Stable Diffusion 2.0 to achieve precise image generation, segmentation, and object detection.

  • Updated May 13, 2024
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the vision-models topic, visit your repo's landing page and select "manage topics."

Learn more