Instructions to use jinaai/ReaderLM-v2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jinaai/ReaderLM-v2 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="jinaai/ReaderLM-v2") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("jinaai/ReaderLM-v2") model = AutoModelForCausalLM.from_pretrained("jinaai/ReaderLM-v2", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Inference
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jinaai/ReaderLM-v2 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jinaai/ReaderLM-v2" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jinaai/ReaderLM-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/jinaai/ReaderLM-v2
- SGLang
How to use jinaai/ReaderLM-v2 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jinaai/ReaderLM-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jinaai/ReaderLM-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jinaai/ReaderLM-v2" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jinaai/ReaderLM-v2", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use jinaai/ReaderLM-v2 with Docker Model Runner:
docker model run hf.co/jinaai/ReaderLM-v2
Optimizing Model Inference Speed for Large HTML Inputs
Hello,
I'm currently utilizing the H20 , ReaderLM model with the max_new_tokens parameter set to 4096, adhering to the official guidelines. My typical input consists of HTML documents ranging from 20 to 40 kb. The inference process generally takes about 4 to 5 minutes. Is this duration considered normal? If so, could you recommend strategies to enhance the inference speed?
Thank you for your assistance.
I believe you can do some simple cleaning of HTML input. Basically, it works.
Thank you for your reply. The HTML I input has been thoroughly cleaned, even more so than the official cleaning logic provided. Since I'm working on real-time search and web-based RAG, some web pages, after cleaning, do become that large.
Currently, after replacing Transformer inference with VLLM, the speed has increased to five times faster than before, and in the best cases, even ten times faster. However, some HTMLs cause errors when using VLLM's tokenizer. Although no error is reported, the phenomenon is that tens of thousands of tokens load per second, and the return is empty. The specific reason is still unclear, but at least it's certain that using VLLM inference can significantly improve speed.
Once again, sincere thanks for your reply.
I believe you can do some simple cleaning of HTML input. Basically, it works.
is it possible to train the model to inference multiple tokens in one predict step? instead of just inferece one token each time.
That's a good idea which demonstrate a good boost on decoding stage. Actually, we are investigating some advanced decoding strategy. We hope we can bring a more efficiency model later at some time point.
mark
Currently, after replacing Transformer inference with VLLM, the speed has increased to five times faster than before, and in the best cases, even ten times faster. However, some HTMLs cause errors when using VLLM's tokenizer. Although no error is reported, the phenomenon is that tens of thousands of tokens load per second, and the return is empty. The specific reason is still unclear, but at least it's certain that using VLLM inference can significantly improve speed.
same issue on vllm
cause endless repetitive content output