Skip to content
#

visual-knowledge-base

Here is 1 public repository matching this topic...

Vision Transformer (ViT)-based pipeline for multimodal image knowledge extraction: fine-grained botanical taxonomy, cultural landmark recognition, and semantic object analysis. Combines pretrained ViTs, domain adapters, and generative language models to produce structured annotations, contextual metadata, and adaptive study resources. + RAG support

  • Updated Aug 8, 2026
  • Python

Add this topic to your repo

To associate your repository with the visual-knowledge-base topic, visit your repo's landing page and select "manage topics."

Learn more