Instructions to use NeverSleep/Noromaid-7b-v0.1.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use NeverSleep/Noromaid-7b-v0.1.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="NeverSleep/Noromaid-7b-v0.1.1")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("NeverSleep/Noromaid-7b-v0.1.1") model = AutoModelForCausalLM.from_pretrained("NeverSleep/Noromaid-7b-v0.1.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use NeverSleep/Noromaid-7b-v0.1.1 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "NeverSleep/Noromaid-7b-v0.1.1" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "NeverSleep/Noromaid-7b-v0.1.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/NeverSleep/Noromaid-7b-v0.1.1
- SGLang
How to use NeverSleep/Noromaid-7b-v0.1.1 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "NeverSleep/Noromaid-7b-v0.1.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "NeverSleep/Noromaid-7b-v0.1.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "NeverSleep/Noromaid-7b-v0.1.1" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "NeverSleep/Noromaid-7b-v0.1.1", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use NeverSleep/Noromaid-7b-v0.1.1 with Docker Model Runner:
docker model run hf.co/NeverSleep/Noromaid-7b-v0.1.1
Noromaid built on top of Starling-LM-7B-alpha
Maywell’s PiVoT-0.1-Starling-LM-RP is based on berkeley-nest/Starling-LM-7B-alpha, which is based on OpenChat 3.5, and it does well at RP. I think Starling-LM-7B-alpha would work as a terrific base for Noromaid.
https://huggingface.co/maywell/PiVoT-0.1-Starling-LM-RP
https://huggingface.co/berkeley-nest/Starling-LM-7B-alpha
https://huggingface.co/openchat/openchat_3.5
OPENCHAT 3.5:
🔥 The first 7B model Achieves Comparable Results with ChatGPT (March)! 🔥
🤖 #1 Open-source model on MT-bench scoring 7.81, outperforming 70B models 🤖
OpenChat is an innovative library of open-source language models, fine-tuned with C-RLFT - a strategy inspired by offline reinforcement learning. Our models learn from mixed-quality data without preference labels, delivering exceptional performance on par with ChatGPT, even with a 7B model. Despite our simple approach, we are committed to developing a high-performance, commercially viable, open-source large language model, and we continue to make significant strides toward this vision.
I got notified for this, and it seems like a good opportunity to shill my PivotMaid-Starling-11b, something that might work for your use case. Sorry if that's out of line.
No worries. I’ll give it a try. I’m always looking for new models to play with. Thanks🤘😁
This is not an official finetune so it will loose quality ofc, if you can show us that it is a good model in practise me way consider training on it.
