Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

DavidAU
/
ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x

Text Generation
Transformers
Safetensors
English
ernie4_5_moe
programming
code generation
code
coding
coder
chat
brainstorm 20x
creative
all uses cases
benchmarks
128k context
creative writing
fiction
finetune
thinking
reasoning
conversational
Model card Files Files and versions
xet
Community
1

Instructions to use DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x with libraries, inference providers, notebooks, and local apps. Follow these links to get started.

  • Libraries
  • Transformers

    How to use DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x with Transformers:

    # Use a pipeline as a high-level helper
    from transformers import pipeline
    
    pipe = pipeline("text-generation", model="DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x")
    messages = [
        {"role": "user", "content": "Who are you?"},
    ]
    pipe(messages)
    # Load model directly
    from transformers import AutoTokenizer, AutoModelForCausalLM
    
    tokenizer = AutoTokenizer.from_pretrained("DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x")
    model = AutoModelForCausalLM.from_pretrained("DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x", device_map="auto")
    messages = [
        {"role": "user", "content": "Who are you?"},
    ]
    inputs = tokenizer.apply_chat_template(
    	messages,
    	add_generation_prompt=True,
    	tokenize=True,
    	return_dict=True,
    	return_tensors="pt",
    ).to(model.device)
    
    outputs = model.generate(**inputs, max_new_tokens=40)
    print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:]))
  • Notebooks
  • Google Colab
  • Kaggle
  • Local Apps Settings
  • vLLM

    How to use DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x with vLLM:

    Install from pip and serve model
    # Install vLLM from pip:
    pip install vllm
    # Start the vLLM server:
    vllm serve "DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x"
    # Call the server using curl (OpenAI-compatible API):
    curl -X POST "http://localhost:8000/v1/chat/completions" \
    	-H "Content-Type: application/json" \
    	--data '{
    		"model": "DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x",
    		"messages": [
    			{
    				"role": "user",
    				"content": "What is the capital of France?"
    			}
    		]
    	}'
    Use Docker
    docker model run hf.co/DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x
  • SGLang

    How to use DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x with SGLang:

    Install from pip and serve model
    # Install SGLang from pip:
    pip install sglang
    # Start the SGLang server:
    python3 -m sglang.launch_server \
        --model-path "DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x" \
        --host 0.0.0.0 \
        --port 30000
    # Call the server using curl (OpenAI-compatible API):
    curl -X POST "http://localhost:30000/v1/chat/completions" \
    	-H "Content-Type: application/json" \
    	--data '{
    		"model": "DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x",
    		"messages": [
    			{
    				"role": "user",
    				"content": "What is the capital of France?"
    			}
    		]
    	}'
    Use Docker images
    docker run --gpus all \
        --shm-size 32g \
        -p 30000:30000 \
        -v ~/.cache/huggingface:/root/.cache/huggingface \
        --env "HF_TOKEN=<secret>" \
        --ipc=host \
        lmsysorg/sglang:latest \
        python3 -m sglang.launch_server \
            --model-path "DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x" \
            --host 0.0.0.0 \
            --port 30000
    # Call the server using curl (OpenAI-compatible API):
    curl -X POST "http://localhost:30000/v1/chat/completions" \
    	-H "Content-Type: application/json" \
    	--data '{
    		"model": "DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x",
    		"messages": [
    			{
    				"role": "user",
    				"content": "What is the capital of France?"
    			}
    		]
    	}'
  • Docker Model Runner

    How to use DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x with Docker Model Runner:

    docker model run hf.co/DavidAU/ERNIE-4.5-37B-A3B-Thinking-Brainstorm20x
New discussion
Resources
  • PR & discussions documentation
  • Code of Conduct
  • Hub documentation

Interested in your fine-tuning work — possible collaboration?

👍 1
1
#1 opened 9 months ago by
ethan7zhanghx
Company
TOS Privacy About Careers
Website
Models Datasets Spaces Pricing Docs