mirror of
https://github.com/DrHo1y/ezrknn-llm.git
synced 2026-10-03 16:46:37 +07:00
release v1.1.1
This commit is contained in:
77
README.md
77
README.md
@@ -1,4 +1,5 @@
|
||||
# Description
|
||||
|
||||
RKLLM software stack can help users to quickly deploy AI models to Rockchip chips. The overall framework is as follows:
|
||||
<center class="half">
|
||||
<div style="background-color:#ffffff;">
|
||||
@@ -14,33 +15,87 @@
|
||||
- RKNPU kernel driver is responsible for interacting with NPU hardware. It has been open source and can be found in the Rockchip kernel code.
|
||||
|
||||
# Support Platform
|
||||
- RK3588 Series
|
||||
- RK3576 Series
|
||||
|
||||
- RK3588 Series
|
||||
- RK3576 Series
|
||||
|
||||
# Support Models
|
||||
- [X] [LLAMA models](https://huggingface.co/meta-llama)
|
||||
- [X] [TinyLLAMA models](https://huggingface.co/TinyLlama)
|
||||
- [X] [Qwen models](https://huggingface.co/models?search=Qwen/Qwen)
|
||||
- [X] [Phi models](https://huggingface.co/models?search=microsoft/phi)
|
||||
- [X] [ChatGLM3-6B](https://huggingface.co/THUDM/chatglm3-6b/tree/103caa40027ebfd8450289ca2f278eac4ff26405)
|
||||
- [X] [Gemma models](https://huggingface.co/collections/google/gemma-2-release-667d6600fd5220e7b967f315)
|
||||
- [X] [InternLM2 models](https://huggingface.co/collections/internlm/internlm2-65b0ce04970888799707893c)
|
||||
- [X] [MiniCPM models](https://huggingface.co/collections/openbmb/minicpm-65d48bf958302b9fd25b698f)
|
||||
|
||||
- [x] [LLAMA models](https://huggingface.co/meta-llama)
|
||||
- [x] [TinyLLAMA models](https://huggingface.co/TinyLlama)
|
||||
- [x] [Qwen models](https://huggingface.co/models?search=Qwen/Qwen)
|
||||
- [x] [Phi models](https://huggingface.co/models?search=microsoft/phi)
|
||||
- [x] [ChatGLM3-6B](https://huggingface.co/THUDM/chatglm3-6b/tree/103caa40027ebfd8450289ca2f278eac4ff26405)
|
||||
- [x] [Gemma models](https://huggingface.co/collections/google/gemma-2-release-667d6600fd5220e7b967f315)
|
||||
- [x] [InternLM2 models](https://huggingface.co/collections/internlm/internlm2-65b0ce04970888799707893c)
|
||||
- [x] [MiniCPM models](https://huggingface.co/collections/openbmb/minicpm-65d48bf958302b9fd25b698f)
|
||||
|
||||
# Model Performance Benchmark
|
||||
|
||||
| model | dtype | seqlen | max_context | new_tokens | TTFT(ms) | Tokens/s | memory(G) | platform |
|
||||
|:-------------- |:---------- |:------:|:-----------:|:----------:|:--------:|:--------:|:---------:|:--------:|
|
||||
| TinyLLAMA-1.1B | w4a16 | 64 | 320 | 256 | 345.00 | 21.10 | 0.77 | RK3576 |
|
||||
| | w4a16_g128 | 64 | 320 | 256 | 410.00 | 18.50 | 0.8 | RK3576 |
|
||||
| | w8a8 | 64 | 320 | 256 | 140.46 | 24.21 | 1.25 | RK3588 |
|
||||
| | w8a8_g512 | 64 | 320 | 256 | 195.00 | 20.08 | 1.29 | RK3588 |
|
||||
| Qwen2-1.5B | w4a16 | 64 | 320 | 256 | 512.00 | 14.40 | 1.75 | RK3576 |
|
||||
| | w4a16_g128 | 64 | 320 | 256 | 550.00 | 12.75 | 1.76 | RK3576 |
|
||||
| | w8a8 | 64 | 320 | 256 | 206.00 | 16.46 | 2.47 | RK3588 |
|
||||
| | w8a8_g128 | 64 | 320 | 256 | 725.00 | 7.00 | 2.65 | RK3588 |
|
||||
| Phi-3-3.8B | w4a16 | 64 | 320 | 256 | 975.00 | 6.60 | 2.16 | RK3576 |
|
||||
| | w4a16_g128 | 64 | 320 | 256 | 1180.00 | 5.85 | 2.23 | RK3576 |
|
||||
| | w8a8 | 64 | 320 | 256 | 516.00 | 7.44 | 3.88 | RK3588 |
|
||||
| | w8a8_g512 | 64 | 320 | 256 | 610.00 | 6.13 | 3.95 | RK3588 |
|
||||
| ChatGLM3-6B | w4a16 | 64 | 320 | 256 | 1168.00 | 4.62 | 3.86 | RK3576 |
|
||||
| | w4a16_g128 | 64 | 320 | 256 | 1582.56 | 3.82 | 3.96 | RK3576 |
|
||||
| | w8a8 | 64 | 320 | 256 | 800.00 | 4.95 | 6.69 | RK3588 |
|
||||
| | w8a8_g128 | 64 | 320 | 256 | 2190.00 | 2.70 | 7.18 | RK3588 |
|
||||
| Gemma2-2B | w4a16 | 64 | 320 | 256 | 628.00 | 8.00 | 3.63 | RK3576 |
|
||||
| | w4a16_g128 | 64 | 320 | 256 | 776.20 | 7.40 | 3.63 | RK3576 |
|
||||
| | w8a8 | 64 | 320 | 256 | 342.29 | 9.67 | 4.84 | RK3588 |
|
||||
| | w8a8_g128 | 64 | 320 | 256 | 1055.00 | 5.49 | 5.14 | RK3588 |
|
||||
| InternLM2-1.8B | w4a16 | 64 | 320 | 256 | 475.00 | 13.30 | 1.59 | RK3576 |
|
||||
| | w4a16_g128 | 64 | 320 | 256 | 572.00 | 11.95 | 1.62 | RK3576 |
|
||||
| | w8a8 | 64 | 320 | 256 | 205.97 | 15.66 | 2.38 | RK3588 |
|
||||
| | w8a8_g512 | 64 | 320 | 256 | 298.00 | 12.66 | 2.45 | RK3588 |
|
||||
| MiniCPM3-4B | w4a16 | 64 | 320 | 256 | 1397.00 | 4.80 | 2.7 | RK3576 |
|
||||
| | w4a16_g128 | 64 | 320 | 256 | 1645.00 | 4.39 | 2.8 | RK3576 |
|
||||
| | w8a8 | 64 | 320 | 256 | 702.18 | 6.15 | 4.65 | RK3588 |
|
||||
| | w8a8_g128 | 64 | 320 | 256 | 1691.00 | 3.42 | 5.06 | RK3588 |
|
||||
| llama3-8B | w4a16 | 64 | 320 | 256 | 1607.98 | 3.60 | 5.63 | RK3576 |
|
||||
| | w4a16_g128 | 64 | 320 | 256 | 2010.00 | 3.00 | 5.76 | RK3576 |
|
||||
| | w8a8 | 64 | 320 | 256 | 1128.00 | 3.79 | 9.21 | RK3588 |
|
||||
| | w8a8_g512 | 64 | 320 | 256 | 1281.35 | 3.05 | 9.45 | RK3588 |
|
||||
|
||||
- This performance data were collected based on the maximum CPU and NPU frequencies of each platform with version 1.1.0.
|
||||
- The script for setting the frequencies is located in the scripts directory.
|
||||
|
||||
# Download
|
||||
|
||||
You can download the latest package, docker image, example, documentation, and platform-tool from [RKLLM_SDK](https://console.zbox.filez.com/l/RJJDmB), fetch code: rkllm
|
||||
|
||||
# Note
|
||||
|
||||
The modifications in version 1.1.0 are significant, making it incompatible with older version models. Please use the latest toolchain for model conversion and inference.
|
||||
- The modifications in version 1.1 are significant, making it incompatible with older version models. Please use the latest toolchain for model conversion and inference.
|
||||
|
||||
- The supported Python versions are:
|
||||
|
||||
- Python 3.8
|
||||
|
||||
- Python 3.10
|
||||
|
||||
- Latest version: [ <u>v1.1.1](https://github.com/airockchip/rknn-llm/releases/tag/release-v1.1.1)</u>
|
||||
|
||||
# RKNN Toolkit2
|
||||
|
||||
If you want to deploy additional AI model, we have introduced a SDK called RKNN-Toolkit2. For details, please refer to:
|
||||
|
||||
https://github.com/airockchip/rknn-toolkit2
|
||||
|
||||
# CHANGELOG
|
||||
|
||||
## v1.1.0
|
||||
|
||||
- Support group-wise quantization (w4a16 group sizes of 32/64/128, w8a8 group sizes of 128/256/512).
|
||||
- Support joint inference with LoRA model loading
|
||||
- Support storage and preloading of prompt cache.
|
||||
|
||||
BIN
doc/Rockchip_RKLLM_SDK_CN_1.1.0.pdf
Executable file → Normal file
BIN
doc/Rockchip_RKLLM_SDK_CN_1.1.0.pdf
Executable file → Normal file
Binary file not shown.
BIN
doc/Rockchip_RKLLM_SDK_EN_1.1.0.pdf
Executable file → Normal file
BIN
doc/Rockchip_RKLLM_SDK_EN_1.1.0.pdf
Executable file → Normal file
Binary file not shown.
@@ -8,8 +8,8 @@ Before running the demo, you need to prepare the following files:
|
||||
### Build
|
||||
You can run the demo with the only command:
|
||||
```bash
|
||||
# Usage: ./build_rkllm_server_flask.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] --npu_num [NPU Core Count] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]
|
||||
./build_rkllm_server_flask.sh --workshop /user/data --model_path /user/data/model.rkllm --platform rk3588 --npu_num 3
|
||||
# Usage: ./build_rkllm_server_flask.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]
|
||||
./build_rkllm_server_flask.sh --workshop /user/data --model_path /user/data/model.rkllm --platform rk3588
|
||||
```
|
||||
### Access with API
|
||||
After building the RKLLM-Server-Flask, You can use ‘chat_api_flask.py’ to access the RKLLM-Server-Flask and get the answser of RKLLM models.
|
||||
@@ -20,8 +20,8 @@ Attention: you should check the IP address of the board with 'ifconfig' command
|
||||
### Build
|
||||
You can run the demo with the only command:
|
||||
```bash
|
||||
# Usage: ./build_rkllm_server_gradio.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] --npu_num [NPU Core Count] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]
|
||||
./build_rkllm_server_gradio.sh --workshop /user/data --model_path /user/data/model.rkllm --platform rk3588 --npu_num 3
|
||||
# Usage: ./build_rkllm_server_gradio.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]
|
||||
./build_rkllm_server_gradio.sh --workshop /user/data --model_path /user/data/model.rkllm --platform rk3588
|
||||
```
|
||||
### Access the Server
|
||||
After running the demo, You can access the RKLLM-Server-Gradio with two ways:
|
||||
|
||||
@@ -3,8 +3,8 @@
|
||||
#*****************************************************************************************#
|
||||
# This script is an automated setup script for the RKLLM-Server-Flask service.
|
||||
# Users can run this script to automate the deployment of the RKLLM-Server-Flask service on a Linux board.
|
||||
# Usage: ./build_rkllm_server_flask.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] --npu_num [NPU Core Count] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]
|
||||
# example: ./build_rkllm_server_flask.sh --workshop /user/data --model_path /user/data/model.rkllm --platform rk3588 --npu_num 3
|
||||
# Usage: ./build_rkllm_server_flask.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]
|
||||
# example: ./build_rkllm_server_flask.sh --workshop /user/data --model_path /user/data/model.rkllm --platform rk3588
|
||||
#*****************************************************************************************#
|
||||
|
||||
LORA_PATH=""
|
||||
@@ -12,7 +12,7 @@ PROMPT_FILE_PATH=""
|
||||
|
||||
# Function to display help
|
||||
function show_help {
|
||||
echo "Usage: ./build_rkllm_server_flask.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] --npu_num [NPU Core Count] [--lora_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]"
|
||||
echo "Usage: ./build_rkllm_server_flask.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] [--lora_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]"
|
||||
}
|
||||
|
||||
# Parse command-line options
|
||||
@@ -30,10 +30,6 @@ while [[ $# -gt 0 ]]; do
|
||||
TARGET_PLATFORM="$2"
|
||||
shift 2
|
||||
;;
|
||||
--npu_num)
|
||||
NPU_CORE_COUNT="$2"
|
||||
shift 2
|
||||
;;
|
||||
--lora_model_path)
|
||||
LORA_PATH="$2"
|
||||
shift 2
|
||||
@@ -93,7 +89,7 @@ cp ../../runtime/Linux/librkllm_api/aarch64/librkllmrt.so ./rkllm_server/lib/
|
||||
adb push ./rkllm_server $WORKING_PATH
|
||||
|
||||
#################### Enter the board terminal and start the server service. ####################
|
||||
CMD="python3 flask_server.py --rkllm_model_path $MODEL_PATH --target_platform $TARGET_PLATFORM --num_npu_core $NPU_CORE_COUNT"
|
||||
CMD="python3 flask_server.py --rkllm_model_path $MODEL_PATH --target_platform $TARGET_PLATFORM"
|
||||
if [[ -n "$LORA_PATH" ]]; then
|
||||
CMD="$CMD --lora_model_path $LORA_PATH"
|
||||
fi
|
||||
|
||||
@@ -3,8 +3,8 @@
|
||||
#*****************************************************************************************#
|
||||
# This script is an automated setup script for the RKLLM-Server-Gradio service.
|
||||
# Users can run this script to automate the deployment of the RKLLM-Server-Gradio service on a Linux board.
|
||||
# Usage: ./build_rkllm_server_gradio.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] --npu_num [NPU Core Count] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]
|
||||
# example: ./build_rkllm_server_gradio.sh --workshop /user/data --model_path /user/data/model.rkllm --platform rk3588 --npu_num 3
|
||||
# Usage: ./build_rkllm_server_gradio.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] [--lora_model_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]
|
||||
# example: ./build_rkllm_server_gradio.sh --workshop /user/data --model_path /user/data/model.rkllm --platform rk3588
|
||||
#*****************************************************************************************#
|
||||
|
||||
LORA_PATH=""
|
||||
@@ -12,7 +12,7 @@ PROMPT_FILE_PATH=""
|
||||
|
||||
# Function to display help
|
||||
function show_help {
|
||||
echo "Usage: ./build_rkllm_server_gradio.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] --npu_num [NPU Core Count] [--lora_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]"
|
||||
echo "Usage: ./build_rkllm_server_gradio.sh --workshop [RKLLM-Server Working Path] --model_path [Absolute Path of Converted RKLLM Model on Board] --platform [Target Platform: rk3588/rk3576] [--lora_path [Lora Model Path]] [--prompt_cache_path [Prompt Cache File Path]]"
|
||||
}
|
||||
|
||||
# Parse command-line options
|
||||
@@ -30,10 +30,6 @@ while [[ $# -gt 0 ]]; do
|
||||
TARGET_PLATFORM="$2"
|
||||
shift 2
|
||||
;;
|
||||
--npu_num)
|
||||
NPU_CORE_COUNT="$2"
|
||||
shift 2
|
||||
;;
|
||||
--lora_model_path)
|
||||
LORA_PATH="$2"
|
||||
shift 2
|
||||
@@ -92,7 +88,7 @@ cp ../../runtime/Linux/librkllm_api/aarch64/librkllmrt.so ./rkllm_server/lib/
|
||||
adb push ./rkllm_server $WORKING_PATH
|
||||
|
||||
#################### Enter the board terminal and start the server service. ####################
|
||||
CMD="python3 gradio_server.py --rkllm_model_path $MODEL_PATH --target_platform $TARGET_PLATFORM --num_npu_core $NPU_CORE_COUNT"
|
||||
CMD="python3 gradio_server.py --rkllm_model_path $MODEL_PATH --target_platform $TARGET_PLATFORM"
|
||||
|
||||
if [[ -n "$LORA_PATH" ]]; then
|
||||
CMD="$CMD --lora_model_path $LORA_PATH"
|
||||
|
||||
@@ -23,9 +23,10 @@ userdata = ctypes.c_void_p(None)
|
||||
|
||||
LLMCallState = ctypes.c_int
|
||||
LLMCallState.RKLLM_RUN_NORMAL = 0
|
||||
LLMCallState.RKLLM_RUN_FINISH = 1
|
||||
LLMCallState.RKLLM_RUN_ERROR = 2
|
||||
LLMCallState.RKLLM_RUN_GET_LAST_HIDDEN_LAYER = 3
|
||||
LLMCallState.RKLLM_RUN_WAITING = 1
|
||||
LLMCallState.RKLLM_RUN_FINISH = 2
|
||||
LLMCallState.RKLLM_RUN_ERROR = 3
|
||||
LLMCallState.RKLLM_RUN_GET_LAST_HIDDEN_LAYER = 4
|
||||
|
||||
RKLLMInputMode = ctypes.c_int
|
||||
RKLLMInputMode.RKLLM_INPUT_PROMPT = 0
|
||||
@@ -37,15 +38,17 @@ RKLLMInferMode = ctypes.c_int
|
||||
RKLLMInferMode.RKLLM_INFER_GENERATE = 0
|
||||
RKLLMInferMode.RKLLM_INFER_GET_LAST_HIDDEN_LAYER = 1
|
||||
|
||||
class RKLLMExtendParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("base_domain_id", ctypes.c_int32),
|
||||
("reserved", ctypes.c_uint8 * 112)
|
||||
]
|
||||
|
||||
class RKLLMParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("model_path", ctypes.c_char_p),
|
||||
("license_path", ctypes.c_char_p),
|
||||
("num_npu_core", ctypes.c_int32),
|
||||
("max_context_len", ctypes.c_int32),
|
||||
("n_prefill_batch", ctypes.c_int32),
|
||||
("max_new_tokens", ctypes.c_int32),
|
||||
("skip_special_token", ctypes.c_bool),
|
||||
("top_k", ctypes.c_int32),
|
||||
("top_p", ctypes.c_float),
|
||||
("temperature", ctypes.c_float),
|
||||
@@ -55,12 +58,12 @@ class RKLLMParam(ctypes.Structure):
|
||||
("mirostat", ctypes.c_int32),
|
||||
("mirostat_tau", ctypes.c_float),
|
||||
("mirostat_eta", ctypes.c_float),
|
||||
("logprobs", ctypes.c_bool),
|
||||
("top_logprobs", ctypes.c_int32),
|
||||
("skip_special_token", ctypes.c_bool),
|
||||
("is_async", ctypes.c_bool),
|
||||
("img_start", ctypes.c_char_p),
|
||||
("img_end", ctypes.c_char_p),
|
||||
("img_content", ctypes.c_char_p),
|
||||
("extend_param", RKLLMExtendParam),
|
||||
]
|
||||
|
||||
class RKLLMLoraAdapter(ctypes.Structure):
|
||||
@@ -70,17 +73,6 @@ class RKLLMLoraAdapter(ctypes.Structure):
|
||||
("scale", ctypes.c_float)
|
||||
]
|
||||
|
||||
class RKLLMLoraParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("lora_adapter_name", ctypes.c_char_p)
|
||||
]
|
||||
|
||||
class RKLLMPromptCacheParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("save_prompt_cache", ctypes.c_int),
|
||||
("prompt_cache_path", ctypes.c_char_p)
|
||||
]
|
||||
|
||||
class RKLLMEmbedInput(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("embed", ctypes.POINTER(ctypes.c_float)),
|
||||
@@ -113,6 +105,18 @@ class RKLLMInput(ctypes.Structure):
|
||||
("input_mode", ctypes.c_int),
|
||||
("input_data", RKLLMInputUnion)
|
||||
]
|
||||
|
||||
class RKLLMLoraParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("lora_adapter_name", ctypes.c_char_p)
|
||||
]
|
||||
|
||||
class RKLLMPromptCacheParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("save_prompt_cache", ctypes.c_int),
|
||||
("prompt_cache_path", ctypes.c_char_p)
|
||||
]
|
||||
|
||||
class RKLLMInferParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("mode", RKLLMInferMode),
|
||||
@@ -199,14 +203,11 @@ callback = callback_type(callback_impl)
|
||||
|
||||
# Define the RKLLM class, which includes initialization, inference, and release operations for the RKLLM model in the dynamic library
|
||||
class RKLLM(object):
|
||||
def __init__(self, model_path, num_npu_core, lora_model_path = None, prompt_cache_path = None):
|
||||
def __init__(self, model_path, lora_model_path = None, prompt_cache_path = None):
|
||||
rkllm_param = RKLLMParam()
|
||||
rkllm_param.model_path = bytes(model_path, 'utf-8')
|
||||
rkllm_param.license_path = None
|
||||
rkllm_param.num_npu_core = num_npu_core
|
||||
|
||||
rkllm_param.max_context_len = 512
|
||||
rkllm_param.n_prefill_batch = 512
|
||||
rkllm_param.max_new_tokens = -1
|
||||
rkllm_param.skip_special_token = True
|
||||
|
||||
@@ -221,20 +222,20 @@ class RKLLM(object):
|
||||
rkllm_param.mirostat_tau = 5.0
|
||||
rkllm_param.mirostat_eta = 0.1
|
||||
|
||||
rkllm_param.logprobs = False
|
||||
rkllm_param.top_logprobs = 5
|
||||
rkllm_param.is_async = False
|
||||
|
||||
rkllm_param.img_start = "".encode('utf-8')
|
||||
rkllm_param.img_end = "".encode('utf-8')
|
||||
rkllm_param.img_content = "".encode('utf-8')
|
||||
|
||||
rkllm_param.extend_param.base_domain_id = 0
|
||||
|
||||
self.handle = RKLLM_Handle_t()
|
||||
|
||||
self.rkllm_init = rkllm_lib.rkllm_init
|
||||
self.rkllm_init.argtypes = [ctypes.POINTER(RKLLM_Handle_t), RKLLMParam, callback_type]
|
||||
self.rkllm_init.argtypes = [ctypes.POINTER(RKLLM_Handle_t), ctypes.POINTER(RKLLMParam), callback_type]
|
||||
self.rkllm_init.restype = ctypes.c_int
|
||||
self.rkllm_init(ctypes.byref(self.handle), rkllm_param, callback)
|
||||
self.rkllm_init(ctypes.byref(self.handle), ctypes.byref(rkllm_param), callback)
|
||||
|
||||
self.rkllm_run = rkllm_lib.rkllm_run
|
||||
self.rkllm_run.argtypes = [RKLLM_Handle_t, ctypes.POINTER(RKLLMInput), ctypes.POINTER(RKLLMInferParam), ctypes.c_void_p]
|
||||
@@ -294,7 +295,6 @@ if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument('--rkllm_model_path', type=str, required=True, help='Absolute path of the converted RKLLM model on the Linux board;')
|
||||
parser.add_argument('--target_platform', type=str, required=True, help='Target platform: e.g., rk3588/rk3576;')
|
||||
parser.add_argument('--num_npu_core', type=int, required=True, help='Target num_npu_core;')
|
||||
parser.add_argument('--lora_model_path', type=str, help='Absolute path of the lora_model on the Linux board;')
|
||||
parser.add_argument('--prompt_cache_path', type=str, help='Absolute path of the prompt_cache file on the Linux board;')
|
||||
args = parser.parse_args()
|
||||
@@ -309,11 +309,6 @@ if __name__ == "__main__":
|
||||
sys.stdout.flush()
|
||||
exit()
|
||||
|
||||
if not (args.num_npu_core in [1, 2, 3]):
|
||||
print("Error: rk3576 supports 1/2 cores, rk3588 supports 1/2/3 cores, please specify the correct number of cores.")
|
||||
sys.stdout.flush()
|
||||
exit()
|
||||
|
||||
if args.lora_model_path:
|
||||
if not os.path.exists(args.lora_model_path):
|
||||
print("Error: Please provide the correct lora_model path, and advise it is the absolute path on the board.")
|
||||
@@ -337,8 +332,7 @@ if __name__ == "__main__":
|
||||
print("=========init....===========")
|
||||
sys.stdout.flush()
|
||||
model_path = args.rkllm_model_path
|
||||
num_npu_core = args.num_npu_core
|
||||
rkllm_model = RKLLM(model_path, num_npu_core, args.lora_model_path, args.prompt_cache_path)
|
||||
rkllm_model = RKLLM(model_path, args.lora_model_path, args.prompt_cache_path)
|
||||
print("RKLLM Model has been initialized successfully!")
|
||||
print("==============================")
|
||||
sys.stdout.flush()
|
||||
|
||||
@@ -24,9 +24,10 @@ userdata = ctypes.c_void_p(None)
|
||||
|
||||
LLMCallState = ctypes.c_int
|
||||
LLMCallState.RKLLM_RUN_NORMAL = 0
|
||||
LLMCallState.RKLLM_RUN_FINISH = 1
|
||||
LLMCallState.RKLLM_RUN_ERROR = 2
|
||||
LLMCallState.RKLLM_RUN_GET_LAST_HIDDEN_LAYER = 3
|
||||
LLMCallState.RKLLM_RUN_WAITING = 1
|
||||
LLMCallState.RKLLM_RUN_FINISH = 2
|
||||
LLMCallState.RKLLM_RUN_ERROR = 3
|
||||
LLMCallState.RKLLM_RUN_GET_LAST_HIDDEN_LAYER = 4
|
||||
|
||||
RKLLMInputMode = ctypes.c_int
|
||||
RKLLMInputMode.RKLLM_INPUT_PROMPT = 0
|
||||
@@ -38,15 +39,17 @@ RKLLMInferMode = ctypes.c_int
|
||||
RKLLMInferMode.RKLLM_INFER_GENERATE = 0
|
||||
RKLLMInferMode.RKLLM_INFER_GET_LAST_HIDDEN_LAYER = 1
|
||||
|
||||
class RKLLMExtendParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("base_domain_id", ctypes.c_int32),
|
||||
("reserved", ctypes.c_uint8 * 112)
|
||||
]
|
||||
|
||||
class RKLLMParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("model_path", ctypes.c_char_p),
|
||||
("license_path", ctypes.c_char_p),
|
||||
("num_npu_core", ctypes.c_int32),
|
||||
("max_context_len", ctypes.c_int32),
|
||||
("n_prefill_batch", ctypes.c_int32),
|
||||
("max_new_tokens", ctypes.c_int32),
|
||||
("skip_special_token", ctypes.c_bool),
|
||||
("top_k", ctypes.c_int32),
|
||||
("top_p", ctypes.c_float),
|
||||
("temperature", ctypes.c_float),
|
||||
@@ -56,12 +59,12 @@ class RKLLMParam(ctypes.Structure):
|
||||
("mirostat", ctypes.c_int32),
|
||||
("mirostat_tau", ctypes.c_float),
|
||||
("mirostat_eta", ctypes.c_float),
|
||||
("logprobs", ctypes.c_bool),
|
||||
("top_logprobs", ctypes.c_int32),
|
||||
("skip_special_token", ctypes.c_bool),
|
||||
("is_async", ctypes.c_bool),
|
||||
("img_start", ctypes.c_char_p),
|
||||
("img_end", ctypes.c_char_p),
|
||||
("img_content", ctypes.c_char_p),
|
||||
("extend_param", RKLLMExtendParam),
|
||||
]
|
||||
|
||||
class RKLLMLoraAdapter(ctypes.Structure):
|
||||
@@ -71,17 +74,6 @@ class RKLLMLoraAdapter(ctypes.Structure):
|
||||
("scale", ctypes.c_float)
|
||||
]
|
||||
|
||||
class RKLLMLoraParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("lora_adapter_name", ctypes.c_char_p)
|
||||
]
|
||||
|
||||
class RKLLMPromptCacheParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("save_prompt_cache", ctypes.c_int),
|
||||
("prompt_cache_path", ctypes.c_char_p)
|
||||
]
|
||||
|
||||
class RKLLMEmbedInput(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("embed", ctypes.POINTER(ctypes.c_float)),
|
||||
@@ -114,6 +106,18 @@ class RKLLMInput(ctypes.Structure):
|
||||
("input_mode", ctypes.c_int),
|
||||
("input_data", RKLLMInputUnion)
|
||||
]
|
||||
|
||||
class RKLLMLoraParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("lora_adapter_name", ctypes.c_char_p)
|
||||
]
|
||||
|
||||
class RKLLMPromptCacheParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("save_prompt_cache", ctypes.c_int),
|
||||
("prompt_cache_path", ctypes.c_char_p)
|
||||
]
|
||||
|
||||
class RKLLMInferParam(ctypes.Structure):
|
||||
_fields_ = [
|
||||
("mode", RKLLMInferMode),
|
||||
@@ -193,14 +197,11 @@ callback = callback_type(callback_impl)
|
||||
|
||||
# Define the RKLLM class, which includes initialization, inference, and release operations for the RKLLM model in the dynamic library
|
||||
class RKLLM(object):
|
||||
def __init__(self, model_path, num_npu_core, lora_model_path = None, prompt_cache_path = None):
|
||||
def __init__(self, model_path, lora_model_path = None, prompt_cache_path = None):
|
||||
rkllm_param = RKLLMParam()
|
||||
rkllm_param.model_path = bytes(model_path, 'utf-8')
|
||||
rkllm_param.license_path = None
|
||||
rkllm_param.num_npu_core = num_npu_core
|
||||
|
||||
rkllm_param.max_context_len = 512
|
||||
rkllm_param.n_prefill_batch = 512
|
||||
rkllm_param.max_new_tokens = -1
|
||||
rkllm_param.skip_special_token = True
|
||||
|
||||
@@ -215,20 +216,20 @@ class RKLLM(object):
|
||||
rkllm_param.mirostat_tau = 5.0
|
||||
rkllm_param.mirostat_eta = 0.1
|
||||
|
||||
rkllm_param.logprobs = False
|
||||
rkllm_param.top_logprobs = 5
|
||||
rkllm_param.is_async = False
|
||||
|
||||
rkllm_param.img_start = "".encode('utf-8')
|
||||
rkllm_param.img_end = "".encode('utf-8')
|
||||
rkllm_param.img_content = "".encode('utf-8')
|
||||
|
||||
rkllm_param.extend_param.base_domain_id = 0
|
||||
|
||||
self.handle = RKLLM_Handle_t()
|
||||
|
||||
self.rkllm_init = rkllm_lib.rkllm_init
|
||||
self.rkllm_init.argtypes = [ctypes.POINTER(RKLLM_Handle_t), RKLLMParam, callback_type]
|
||||
self.rkllm_init.argtypes = [ctypes.POINTER(RKLLM_Handle_t), ctypes.POINTER(RKLLMParam), callback_type]
|
||||
self.rkllm_init.restype = ctypes.c_int
|
||||
self.rkllm_init(ctypes.byref(self.handle), rkllm_param, callback)
|
||||
self.rkllm_init(ctypes.byref(self.handle), ctypes.byref(rkllm_param), callback)
|
||||
|
||||
self.rkllm_run = rkllm_lib.rkllm_run
|
||||
self.rkllm_run.argtypes = [RKLLM_Handle_t, ctypes.POINTER(RKLLMInput), ctypes.POINTER(RKLLMInferParam), ctypes.c_void_p]
|
||||
@@ -288,7 +289,6 @@ if __name__ == "__main__":
|
||||
parser = argparse.ArgumentParser()
|
||||
parser.add_argument('--rkllm_model_path', type=str, required=True, help='Absolute path of the converted RKLLM model on the Linux board;')
|
||||
parser.add_argument('--target_platform', type=str, required=True, help='Target platform: e.g., rk3588/rk3576;')
|
||||
parser.add_argument('--num_npu_core', type=int, required=True, help='Target num_npu_core;')
|
||||
parser.add_argument('--lora_model_path', type=str, help='Absolute path of the lora_model on the Linux board;')
|
||||
parser.add_argument('--prompt_cache_path', type=str, help='Absolute path of the prompt_cache file on the Linux board;')
|
||||
args = parser.parse_args()
|
||||
@@ -303,11 +303,6 @@ if __name__ == "__main__":
|
||||
sys.stdout.flush()
|
||||
exit()
|
||||
|
||||
if not (args.num_npu_core in [1, 2, 3]):
|
||||
print("Error: rk3576 supports 1/2 cores, rk3588 supports 1/2/3 cores, please specify the correct number of cores.")
|
||||
sys.stdout.flush()
|
||||
exit()
|
||||
|
||||
if args.lora_model_path:
|
||||
if not os.path.exists(args.lora_model_path):
|
||||
print("Error: Please provide the correct lora_model path, and advise it is the absolute path on the board.")
|
||||
@@ -331,8 +326,7 @@ if __name__ == "__main__":
|
||||
print("=========init....===========")
|
||||
sys.stdout.flush()
|
||||
model_path = args.rkllm_model_path
|
||||
num_npu_core = args.num_npu_core
|
||||
rkllm_model = RKLLM(model_path, num_npu_core, args.lora_model_path, args.prompt_cache_path)
|
||||
rkllm_model = RKLLM(model_path, args.lora_model_path, args.prompt_cache_path)
|
||||
print("RKLLM Model has been initialized successfully!")
|
||||
print("==============================")
|
||||
sys.stdout.flush()
|
||||
|
||||
Binary file not shown.
Binary file not shown.
@@ -1 +1,2 @@
|
||||
1e69c83aa9b718a5b75e368aa51d0ac4 rkllm_toolkit-1.1.0-cp38-cp38-linux_x86_64.whl
|
||||
696070d1ec134ff68fd5b03b873cd678 rkllm_toolkit-1.1.1-cp38-cp38-linux_x86_64.whl
|
||||
e8e37c6bf0af799f9a0187c9864b2450 rkllm_toolkit-1.1.1-cp310-cp310-linux_x86_64.whl
|
||||
|
||||
Binary file not shown.
Binary file not shown.
43
scripts/fix_freq_rk3576.sh
Executable file
43
scripts/fix_freq_rk3576.sh
Executable file
@@ -0,0 +1,43 @@
|
||||
#!/system/bin/sh
|
||||
|
||||
echo 1 > /sys/devices/system/cpu/cpu0/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu1/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu2/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu3/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu4/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu5/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu6/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu7/cpuidle/state1/disable
|
||||
|
||||
echo "NPU available frequencies:"
|
||||
cat /sys/class/devfreq/27700000.npu/available_frequencies
|
||||
echo "Fix NPU max frequency:"
|
||||
echo userspace > /sys/class/devfreq/27700000.npu/governor
|
||||
echo 1000000000 > /sys/class/devfreq/27700000.npu/userspace/set_freq
|
||||
cat /sys/class/devfreq/27700000.npu/cur_freq
|
||||
|
||||
echo "CPU available frequencies:"
|
||||
cat /sys/devices/system/cpu/cpufreq/policy0/scaling_available_frequencies
|
||||
cat /sys/devices/system/cpu/cpufreq/policy4/scaling_available_frequencies
|
||||
echo "Fix CPU max frequency:"
|
||||
echo userspace > /sys/devices/system/cpu/cpufreq/policy0/scaling_governor
|
||||
echo 2208000 > /sys/devices/system/cpu/cpufreq/policy0/scaling_setspeed
|
||||
cat /sys/devices/system/cpu/cpufreq/policy0/scaling_cur_freq
|
||||
echo userspace > /sys/devices/system/cpu/cpufreq/policy4/scaling_governor
|
||||
echo 2304000 > /sys/devices/system/cpu/cpufreq/policy4/scaling_setspeed
|
||||
cat /sys/devices/system/cpu/cpufreq/policy4/scaling_cur_freq
|
||||
|
||||
echo "GPU available frequencies:"
|
||||
cat /sys/class/devfreq/27800000.gpu/cur_freq
|
||||
cat /sys/class/devfreq/27800000.gpu/available_frequencies
|
||||
echo "Fix GPU max frequency:"
|
||||
echo userspace > /sys/class/devfreq/27800000.gpu/governor
|
||||
echo 950000000 > /sys/class/devfreq/27800000.gpu/userspace/set_freq
|
||||
cat /sys/class/devfreq/27800000.gpu/cur_freq
|
||||
|
||||
echo "DDR available frequencies:"
|
||||
cat /sys/class/devfreq/dmc/available_frequencies
|
||||
echo "Fix DDR max frequency:"
|
||||
echo userspace > /sys/class/devfreq/dmc/governor
|
||||
echo 2112000000 > /sys/class/devfreq/dmc/userspace/set_freq
|
||||
cat /sys/class/devfreq/dmc/cur_freq
|
||||
46
scripts/fix_freq_rk3588.sh
Executable file
46
scripts/fix_freq_rk3588.sh
Executable file
@@ -0,0 +1,46 @@
|
||||
#!/system/bin/sh
|
||||
|
||||
echo 1 > /sys/devices/system/cpu/cpu0/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu1/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu2/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu3/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu4/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu5/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu6/cpuidle/state1/disable
|
||||
echo 1 > /sys/devices/system/cpu/cpu7/cpuidle/state1/disable
|
||||
|
||||
echo "NPU available frequencies:"
|
||||
cat /sys/class/devfreq/fdab0000.npu/available_frequencies
|
||||
echo "Fix NPU max frequency:"
|
||||
echo userspace > /sys/class/devfreq/fdab0000.npu/governor
|
||||
echo 1000000000 > /sys/class/devfreq/fdab0000.npu/userspace/set_freq
|
||||
cat /sys/class/devfreq/fdab0000.npu/cur_freq
|
||||
|
||||
echo "CPU available frequencies:"
|
||||
cat /sys/devices/system/cpu/cpufreq/policy0/scaling_available_frequencies
|
||||
cat /sys/devices/system/cpu/cpufreq/policy4/scaling_available_frequencies
|
||||
cat /sys/devices/system/cpu/cpufreq/policy6/scaling_available_frequencies
|
||||
echo "Fix CPU max frequency:"
|
||||
echo userspace > /sys/devices/system/cpu/cpufreq/policy0/scaling_governor
|
||||
echo 1800000 > /sys/devices/system/cpu/cpufreq/policy0/scaling_setspeed
|
||||
cat /sys/devices/system/cpu/cpufreq/policy0/scaling_cur_freq
|
||||
echo userspace > /sys/devices/system/cpu/cpufreq/policy4/scaling_governor
|
||||
echo 2352000 > /sys/devices/system/cpu/cpufreq/policy4/scaling_setspeed
|
||||
cat /sys/devices/system/cpu/cpufreq/policy4/scaling_cur_freq
|
||||
echo userspace > /sys/devices/system/cpu/cpufreq/policy6/scaling_governor
|
||||
echo 2352000 > /sys/devices/system/cpu/cpufreq/policy6/scaling_setspeed
|
||||
cat /sys/devices/system/cpu/cpufreq/policy6/scaling_cur_freq
|
||||
|
||||
echo "GPU available frequencies:"
|
||||
cat /sys/class/devfreq/fb000000.gpu/available_frequencies
|
||||
echo "Fix GPU max frequency:"
|
||||
echo userspace > /sys/class/devfreq/fb000000.gpu/governor
|
||||
echo 1000000000 > /sys/class/devfreq/fb000000.gpu/userspace/set_freq
|
||||
cat /sys/class/devfreq/fb000000.gpu/cur_freq
|
||||
|
||||
echo "DDR available frequencies:"
|
||||
cat /sys/class/devfreq/dmc/available_frequencies
|
||||
echo "Fix DDR max frequency:"
|
||||
echo userspace > /sys/class/devfreq/dmc/governor
|
||||
echo 2112000000 > /sys/class/devfreq/dmc/userspace/set_freq
|
||||
cat /sys/class/devfreq/dmc/cur_freq
|
||||
Reference in New Issue
Block a user