A collection of Jupyter notebooks showcasing the use of Generative AI models, including Large Language Models (LLMs), Vision-Language Models (VLMs), and Diffusion Models
Install uv, then use the root project for repository tooling:
uv sync
uv run python tools/validate_metadata.py
uv run python tools/update_readme.pyFor notebook execution, use the environment shown in the matrix below. Shared environments live under envs/; models marked model-specific need the setup notes in their model folder metadata.
Register a shared environment as a Jupyter kernel:
cd envs/hf-transformers
uv sync
uv run python -m ipykernel install --user --name genai-hf-transformers- AEMatter — segmentation-matting; matting, segmentation
- Alfie — generation; image-generation, rgba-generation
- AnimateDiff — generation; video-generation, diffusion
- AnomalyCLIP — anomaly; anomaly-detection, zero-shot
- AuraSR — generation; super-resolution, image-generation
- BART — llm; text-generation, seq2seq
- BEN2 — segmentation-matting; segmentation, HR-seg, DIS
- BERT-deepset — llm; question-answering
- BiRefNet — segmentation-matting; matting, segmentation, DIS
- BLIP — vlm; VLM, captioning, VQA, image-text-retrieval
- CAT-Seg — segmentation-matting; segmentation, open-vocabulary-segmentation
- CLIP — retrieval-embedding; contrastive, image-features, text-features, retrieval
- CLIPseg — segmentation-matting; segmentation, text-prompted-segmentation
- CogVLM — vlm; VLM, VQA
- ControlNet — generation; image-generation, conditioned-generation, diffusion
- DepthAnything — depth-3d; depth-estimation
- DepthPro — depth-3d; depth-estimation, metric-depth
- DETR — detection; detection
- DiffDIS — segmentation-matting; segmentation, HR-seg, DIS
- DINOv3 — classification-representation; image-features, representation-learning
- DPT — depth-3d; depth-estimation
- EoMT — classification-representation; segmentation, image-features
- EssentialAI — llm; text-generation
- EVA — classification-representation; image-classification, image-features
- FaceParsing — segmentation-matting; segmentation, face-parsing
- Ferret — vlm; VLM, grounded-chat
- FineGrain — segmentation-matting; segmentation, box-prompted-segmentation
- FLAVA — retrieval-embedding; multimodal, retrieval, image-text
- FLUX — generation; image-generation, diffusion
- GCL — retrieval-embedding; retrieval, ranking, contrastive
- Gemma3 — llm; text-generation, VLM
- GLIDE — generation; image-generation, diffusion
- GLIP — detection; detection, grounding
- GroundingDINO — detection; detection, grounding, open-vocabulary-detection
- ImageGPT — generation; image-generation, pixel-modeling
- InternVL — vlm; VLM, VQA
- Janus — vlm; VLM, multimodal-generation
- Leffa — generation; person-image-generation, virtual-try-on
- LeViT — classification-representation; image-classification
- LISA — segmentation-matting; segmentation, reasoning-segmentation, VLM
- LLaMA2 — llm; text-generation, chat
- LLaVA — vlm; VLM, VQA
- LLaVA-NeXT — vlm; VLM, VQA
- LLaVA-OneVision — vlm; VLM, VQA
- Mask2Former — segmentation-matting; segmentation, panoptic-segmentation
- OneFormer — segmentation-matting; segmentation, semantic-segmentation, instance-segmentation, panoptic-segmentation
- OV-DINO — detection; detection, open-vocabulary-detection
- OVSeg — segmentation-matting; segmentation, open-vocabulary-segmentation
- OWL-v2 — detection; detection, open-vocabulary-detection
- OWL-ViT (owlvit_huggingface_inference.ipynb) — detection; detection, open-vocabulary-detection
- OWL-ViT (owlvit_inference-2.ipynb) — detection; detection, open-vocabulary-detection
- PoolFormer — classification-representation; image-classification, image-features
- PromptDepthAnything — depth-3d; depth-estimation, prompted-depth
- QLIP — retrieval-embedding; image-features, text-aligned-tokenization
- SA2VA — vlm; VLM, segmentation, grounded-understanding
- SAM — segmentation-matting; segmentation, prompted-segmentation
- SAM-2 — segmentation-matting; segmentation, prompted-segmentation
- SAM-3 — segmentation-matting; segmentation, concept-prompted-segmentation
- SAM-HQ — segmentation-matting; segmentation, high-quality-segmentation
- SAMRefiner — segmentation-matting; segmentation, seg-refinement
- SAN — segmentation-matting; segmentation, open-vocabulary-segmentation
- SD2 — generation; image-generation, diffusion
- SegFormer — segmentation-matting; segmentation, clothes-segmentation
- SegZero — segmentation-matting; segmentation, reasoning-segmentation
- SigLIP — retrieval-embedding; contrastive, image-features, text-features, retrieval
- SmolVLM — vlm; VLM, VQA
- UNO — generation; image-generation, in-context-generation
- UperNet — segmentation-matting; segmentation, scene-understanding
- VGGT — depth-3d; 3D, visual-geometry, depth-estimation
- VisionReasoner — vlm; VLM, visual-reasoning
- WebSSL — retrieval-embedding; image-features, similarity
- xLAM — llm; text-generation, agent-actions
- YOLO-World — detection; detection, open-vocabulary-detection
- YOLOS4Fashion — detection; detection, fashion-detection
anomaly
classification-representation
depth-3d
detection
- DETR
- GLIP
- GroundingDINO
- OV-DINO
- OWL-v2
- OWL-ViT (owlvit_huggingface_inference.ipynb)
- OWL-ViT (owlvit_inference-2.ipynb)
- YOLO-World
- YOLOS4Fashion
generation
llm
retrieval-embedding
segmentation-matting
- AEMatter
- BEN2
- BiRefNet
- CAT-Seg
- CLIPseg
- DiffDIS
- FaceParsing
- FineGrain
- LISA
- Mask2Former
- OneFormer
- OVSeg
- SAM
- SAM-2
- SAM-3
- SAM-HQ
- SAMRefiner
- SAN
- SegFormer
- SegZero
- UperNet
vlm
- 3D: VGGT
- agent-actions: xLAM
- anomaly-detection: AnomalyCLIP
- box-prompted-segmentation: FineGrain
- captioning: BLIP
- chat: LLaMA2
- clothes-segmentation: SegFormer
- concept-prompted-segmentation: SAM-3
- conditioned-generation: ControlNet
- contrastive: CLIP, GCL, SigLIP
- depth-estimation: DepthAnything, DepthPro, DPT, PromptDepthAnything, VGGT
- detection: DETR, GLIP, GroundingDINO, OV-DINO, OWL-v2, OWL-ViT (owlvit_huggingface_inference.ipynb), OWL-ViT (owlvit_inference-2.ipynb), YOLO-World, YOLOS4Fashion
- diffusion: AnimateDiff, ControlNet, FLUX, GLIDE, SD2
- DIS: BEN2, BiRefNet, DiffDIS
- face-parsing: FaceParsing
- fashion-detection: YOLOS4Fashion
- grounded-chat: Ferret
- grounded-understanding: SA2VA
- grounding: GLIP, GroundingDINO
- high-quality-segmentation: SAM-HQ
- HR-seg: BEN2, DiffDIS
- image-classification: EVA, LeViT, PoolFormer
- image-features: CLIP, DINOv3, EoMT, EVA, PoolFormer, QLIP, SigLIP, WebSSL
- image-generation: Alfie, AuraSR, ControlNet, FLUX, GLIDE, ImageGPT, SD2, UNO
- image-text: FLAVA
- image-text-retrieval: BLIP
- in-context-generation: UNO
- instance-segmentation: OneFormer
- matting: AEMatter, BiRefNet
- metric-depth: DepthPro
- multimodal: FLAVA
- multimodal-generation: Janus
- open-vocabulary-detection: GroundingDINO, OV-DINO, OWL-v2, OWL-ViT (owlvit_huggingface_inference.ipynb), OWL-ViT (owlvit_inference-2.ipynb), YOLO-World
- open-vocabulary-segmentation: CAT-Seg, OVSeg, SAN
- panoptic-segmentation: Mask2Former, OneFormer
- person-image-generation: Leffa
- pixel-modeling: ImageGPT
- prompted-depth: PromptDepthAnything
- prompted-segmentation: SAM, SAM-2
- question-answering: BERT-deepset
- ranking: GCL
- reasoning-segmentation: LISA, SegZero
- representation-learning: DINOv3
- retrieval: CLIP, FLAVA, GCL, SigLIP
- rgba-generation: Alfie
- scene-understanding: UperNet
- seg-refinement: SAMRefiner
- segmentation: AEMatter, BEN2, BiRefNet, CAT-Seg, CLIPseg, DiffDIS, EoMT, FaceParsing, FineGrain, LISA, Mask2Former, OneFormer, OVSeg, SA2VA, SAM, SAM-2, SAM-3, SAM-HQ, SAMRefiner, SAN, SegFormer, SegZero, UperNet
- semantic-segmentation: OneFormer
- seq2seq: BART
- similarity: WebSSL
- super-resolution: AuraSR
- text-aligned-tokenization: QLIP
- text-features: CLIP, SigLIP
- text-generation: BART, EssentialAI, Gemma3, LLaMA2, xLAM
- text-prompted-segmentation: CLIPseg
- video-generation: AnimateDiff
- virtual-try-on: Leffa
- visual-geometry: VGGT
- visual-reasoning: VisionReasoner
- VLM: BLIP, CogVLM, Ferret, Gemma3, InternVL, Janus, LISA, LLaVA, LLaVA-NeXT, LLaVA-OneVision, SA2VA, SmolVLM, VisionReasoner
- VQA: BLIP, CogVLM, InternVL, LLaVA, LLaVA-NeXT, LLaVA-OneVision, SmolVLM
- zero-shot: AnomalyCLIP
| Model | Notebook | Environment | Python | Notes |
|---|---|---|---|---|
| AEMatter | aematter_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| Alfie | alfie_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| AnimateDiff | animatediff_huggingface_inference.ipynb | diffusers (envs/diffusers) |
>=3.10,<3.13 |
Diffusers-based text/image/video generation notebooks. |
| AnomalyCLIP | anomalyclip_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| AuraSR | aurasr_inference.ipynb | diffusers (envs/diffusers) |
>=3.10,<3.13 |
Diffusers-based text/image/video generation notebooks. |
| BART | bart_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| BEN2 | ben2_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| BERT-deepset | bert-deepset_haystack_inference.ipynb | haystack (envs/haystack) |
>=3.10,<3.13 |
Haystack question-answering notebook. |
| BiRefNet | birefnet-hr-matting_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| BLIP | blip_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| CAT-Seg | catseg_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Requires Detectron2 and project-specific open-vocabulary segmentation dependencies. |
| CLIP | clip_huggingface_inference.ipynb | retrieval (envs/retrieval) |
>=3.10,<3.13 |
Embedding, CLIP/SigLIP, and retrieval notebooks. |
| CLIPseg | clipseg_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| CogVLM | cogvlm_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| ControlNet | controlnet_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| DepthAnything | depthanything_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| DepthPro | depthpro_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| DETR | detr_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| DiffDIS | diffdis_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| DINOv3 | dinov3_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| DPT | dpt_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| EoMT | eomt_huggingface_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| EssentialAI | essentialai_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| EVA | eva_timm_inference.ipynb | timm (envs/timm) |
>=3.10,<3.13 |
timm and lightweight representation model notebooks. |
| FaceParsing | faceparsing_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| Ferret | ferret_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Uses a nested project with pinned dependencies and a git Transformers source. |
| FineGrain | finegrainBoxSeg_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| FLAVA | flava_huggingface_inference.ipynb | retrieval (envs/retrieval) |
>=3.10,<3.13 |
Embedding, CLIP/SigLIP, and retrieval notebooks. |
| FLUX | flux_huggingface_inference.ipynb | diffusers (envs/diffusers) |
>=3.10,<3.13 |
Diffusers-based text/image/video generation notebooks. |
| GCL | gcl-e5_huggingface_inference.ipynb | retrieval (envs/retrieval) |
>=3.10,<3.13 |
Embedding, CLIP/SigLIP, and retrieval notebooks. |
| Gemma3 | gemma3_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| GLIDE | glide_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| GLIP | glip_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook documents a Docker image with CUDA 10.2 and PyTorch 1.9; treat as Docker/model-specific unless ported. |
| GroundingDINO | groundingdino_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| ImageGPT | imagegpt_huggingface_inference.ipynb | diffusers (envs/diffusers) |
>=3.10,<3.13 |
Diffusers-based text/image/video generation notebooks. |
| InternVL | internvl_huggingface_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| Janus | janus_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| Leffa | leffa_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| LeViT | levit_timm_inference.ipynb | timm (envs/timm) |
>=3.10,<3.13 |
timm and lightweight representation model notebooks. |
| LISA | lisa_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| LLaMA2 | llama2_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Requires external Llama model access/download approval. |
| LLaVA | llava_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| LLaVA-NeXT | llavanext_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| LLaVA-OneVision | llavaonevision_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| Mask2Former | mask2former_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| OneFormer | oneformer_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| OV-DINO | ovdino_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Requires Detectron2/detrex stack and CUDA-specific Torch pins. |
| OVSeg | ovseg_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Requires Detectron2 and the facebookresearch/ov-seg project setup. |
| OWL-v2 | owlv2_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| OWL-ViT | owlvit_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| OWL-ViT | owlvit_inference-2.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| PoolFormer | poolformer_huggingface_inference.ipynb | timm (envs/timm) |
>=3.10,<3.13 |
timm and lightweight representation model notebooks. |
| PromptDepthAnything | promptdepthanything_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| QLIP | qlip_huggingface_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| SA2VA | sa2va_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| SAM | sam_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| SAM-2 | sam2_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| SAM-3 | sam3_inference.ipynb | sam3 (envs/sam3) |
>=3.12,<3.13 |
Requires Python 3.12, a CUDA-oriented Torch stack, facebookresearch/sam3 editable install, and either local SAM-3 assets or latest Transformers support. |
| SAM-HQ | samhq_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| SAMRefiner | samrefiner_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| SAN | san_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| SD2 | sd_huggingface_inference.ipynb | diffusers (envs/diffusers) |
>=3.10,<3.13 |
Diffusers-based text/image/video generation notebooks. |
| SegFormer | segformer-clothes_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| SegZero | segzero_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| SigLIP | siglip_huggingface_inference.ipynb | retrieval (envs/retrieval) |
>=3.10,<3.13 |
Embedding, CLIP/SigLIP, and retrieval notebooks. |
| SmolVLM | smolvlm_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| UNO | uno_huggingface_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| UperNet | upernet_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| VGGT | vggt_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| VisionReasoner | visionreasoner_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Notebook requires a project clone, pinned CUDA stack, Docker image, or other custom setup. |
| WebSSL | webssl_huggingface_inference.ipynb | retrieval (envs/retrieval) |
>=3.10,<3.13 |
Embedding, CLIP/SigLIP, and retrieval notebooks. |
| xLAM | xlam_huggingface_inferemce.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
| YOLO-World | yolow_inference.ipynb | model-specific (model folder setup) |
see setup_notes |
Requires MMDetection/MMEngine stack and model checkpoint. |
| YOLOS4Fashion | yolos4fashion_huggingface_inference.ipynb | hf-transformers (envs/hf-transformers) |
>=3.10,<3.13 |
Common Hugging Face Transformers vision, text, and multimodal notebooks. |
Notebook execution results are written to reports/notebook_uv_execution.json by tools/execute_notebooks.py. Source notebook outputs are preserved unless they are explicitly inspected and proven safe to remove.