Deployment Guide: Ornith 35B FP8 E4M3 with MTP on RTX Hardware
Hardware &으로 Requirements
- GPU (RTX Based): Requires more than 80GB VRAM to run full context window (256k).
- Memory Systems: Compatible with GB10 Unified Memory Systems if grafting into target NVFP4 models.
Installation & Launch Guide
- Prepare your environment by setting up an appropriate container running enough capacity for large scale weights (
).vLLM high performance inference container - Download requested weight files from GitHub repository
.https://github.com/kyr0/Ornith-35B-FP8-E4M3-MTP - Use the grafter script provided in the repositoryto graft the MTP model into any targeted way, such as onto certain hardware like Hopper and Ada gen targets via specialized scripts.
- Launch the engine utilizing vLLM specifically configured for this grafted architecture.
Optimization & Performance Tips
- Speedup: Using the customly grafted MTP version provides a speed improvement of approximately 18% compared to non-MTP versions.
- Context Management: Ensure system supports or allocates sufficient resources even though it is capable of handling downsize contexts depending on available memory properly managed through the rsetmpt method mentioned (grafting instructions attacheds_script으로 처리 가능할수도 있음 - note using scrippgett manner per source description mention s việc use my script if GB10 type used).
This setup offers an ableisted Pareto optimal path enough heavy lifting tasks with significant throughput gains by leveraging specialized drafting speedse above standard FP8 implementations.
! DYOR (Do Your Own Research)