LensVLM: Compressing long context as images, expanding only relevant pages
4 hours ago
- Apple/LensVLM-9B can be used via libraries like Transformers, vLLM, and SGLang for multimodal vision-language tasks.
- The model processes compressed images of text and selectively expands relevant pages using learned tools.
- Key usage methods include pipelines, direct model loading, API servers, Docker containerization, and local inference scripts.
- The model alternates Python scripts (e.g., run_demo.py) allow custom document processing with compression options (5x, 10x, 15x).
- License terms specify Apple ML Research Model License for model files and Apple Sample Code License for source code.