Represents the multimodal capabilities of a loaded LLM model.
Whether audio processing is supported
Whether image/vision processing is supported