Model Terminal

LLaVA-Medical-8B CLIP-ViT (Stage 2)

A biomedical vision-language model that combines a CLIP ViT vision encoder with an 8B-parameter language model to answer questions about medical images and support multimodal medical reasoning. This checkpoint represents the second-stage instruction-tuned variant fine-tuned on medical data, following the LLaVA-Med two-stage training paradigm. Source: supabase.

Value score
Not scored
Context
tokens
Max output
tokens
Price

Capability radar

Identity

Openness
Modalities
Release
Knowledge cutoff
Deprecation
API docs

Benchmarks

No benchmark scores yet.

Compare nearby

Pricing

Input / 1M
Output / 1M
Speed