ChatGPT   : , , 
 


   : -,     #1
            AI-,      . -  ,  , ,  ,         -, , ,        ,   .  :      .

    Python  ,   .     0      .        .

   .          .     .            .





ChatGPT   : , , 








       .     820     ,       .    ,           .   .   .

             .       pip install    ,      : -  ,  ,    ,  ,   ,  ,       .

  2.0         PDF     .    :    ,     GLM-4.6V-Flash,       HY-MT1.5-7B.   ,  , ZIP-     .               .

  ,     :    ,  ,  .      ,     .        ,   Docker,     .

     . -       ,     ,   .       .     ,    .      .    ,      ,            .

    ,   ,      .      ,           .

    Python, IT-,    , ,     AI-  ,  ,     .  ,    ,   0         .

     .      :        ,        .         ,      .

 :

 (Qwen 32B)

  (HY-MT1.5 Q5)

 

-, , , 

 : PDFPNGOCR 

  :https://github.com/PawelKaev/local-ai-one-button (https://github.com/PawelKaev/local-ai-one-button)

  :

app.py      Qwen, HY-MT, GLM-OCR

gui.py     

local_model.py      

.gitignore   DLL, , PDF

     .









           .    ,   API  ,     .   .    814            .    ,            models/,   .

        .   ,      ,     .     ,  ,   ,       ChatGPT,     .


 1.     

   ChatGPT, Gemini  Copilot   .      ,         GPT.      ,    .

    .       .   ,      ,      .                 ,      .           ,     .             , ,           ChatGPT.           ,      .

   .     .  ,  ,        .     , ,   ,        .      .        ,       OpenAI.  real-time      ,    ,        .    .

    .     .      ,   ...                .    ,   ,   ,   .           1940-   ,         ,      API.   ,        .

    .      .        ,   .           ,    , ,     Slack,          .  Copilot           ,      .


 2.   :   23 

     .   8       .      .

    .    20 ,        23  .       SSD     ,        . DeepSeek  ,   Mixture of Experts   .          Apple M4 Ultra  Snapdragon X Elite.        .

    8 .   Llama 3.2 Vision  11  .    79      ,   ,  .    -.    :    ,   ,       .      ,    .

    .        .          ,  ,  .           LoRA-.      ,         ,     .       Copilot,     .      ,            .

  .        .    35  NPU-     RTX 4070   510 .          ,     ,     .    ,   .


 3.  :  ,    GPT?

    .       :    .          .   8B-       ,  GPT-4o  Claude Sonnet.   .

           .           .             .         API    .       ,      .

   8         ChatGPT    :  , ,  , .      .      ,        ,     ,        .


 

    :

      Qwen 3 32B.

       HY-MT1.5-7B Q5.

      ( , , ).

          .

     :  ,  ,  .

      (  ).

        .

        (   2.0).


   2.0:   

       ,    :

1. PDF  PNG         pdf2image  Poppler.

2. OCR ( )          GLM-4.6V-Flash (GLM-OCR)  Ollama API.

3.          .

4.             HY-MT1.5-7B Q5.

 :

    PDF, ZIP  HTML 

   (  )

      

 ZIP-    PDF

      

      20.2.


,    2.0

:

     HY-MT1.5-7B Q5_K_M (4.6 )     ,   Q4 .     ,    .

    GLM-4.6V-Flash (Q8_0)     .   Ollama API,  ,    .

 :

        .

    ( )    .

        .

          .

 :

     PDF, ZIP, HTML  TXT .

   PDF  ZIP-.

  HTML  TXT  ,  .

    (PNG,  ).

  Ollama:

    GLM-OCR  Ollama API.

        .

:

   20.2       .

       .

      .


 

     :

 app.py   ,       GUI.

 gui.py     Tkinter (8 ).

 local_model.py  -      llama-cpp-python.

  :

 pdf_converter.py   PDF  PNG.

 ocr_extract.py     .

 merge_all_texts.py    .

 translate_with_hy_mt.py   .

 search_and_download.py     .


 

 Python 3.10+

 Ollama (  GLM-OCR)

 Poppler (  PDF)

    CUDA ()  16+  RAM  CPU-

    : ~40   


   

 ,      Python: ,     ,     pip       .       0        .

  ,      -, IT-,       , ,       , ,      -,  ,     AI-   ,        .


 

   ,      ,   .            ,   ,       AI-.

    ,           ,   1  CRM, -  ,                .           .         -     .     ,  ,       ,  .     .


 

,     Python 3.10  .    NVIDIA  6       ,            .        .

?          . !

 0.  ,   .

 0.  ,   :     5 .

    . ,   .   ,       -,    ,           ,   ,      .      .

 ,    ,      .          .            .      .

0.1.   .

   Windows 10  11 (  Mac  Linux  ,    Windows).

  16    ( ,  ).

  25     .

   NVIDIA  6+   (    ,  ).

      .      .

0.2.  1:    -.

   ,     ,  Python,    .     .

1.      install_all.bat    (   ).

2.        ,  C:\LocalAI.

3.       install_all.bat.

4.     .        .   Python,   .         20   .

5.        !  LocalAI.exe   ,    2.

0.3.  2:  .

      -.    .    : , , , .       .

0.4.       .

   ( )        .  .         ChatGPT,   .

    .   PDF, Word      C:\Projects\local-ai-one-button\sample_docs.                 .      .

  ( ).     ,       .        .

   ( ).   ,     (,  , )   .      .

    ( ).  ,  -  ,   .     .

 .    ,         .   .

0.5.  -   .

   ?  C:\Projects\local-ai-one-button\healthcheck.exe ( python healthcheck.py).  ,   .

   ? ,       .   .

   ? ,           Windows.

    ?  -.         ,  .

0.6.  .

   ,    ,               .    1,        ,         ,     AI-.

    .      .




 1. :        





 1.1.    


   ChatGPT   .          .       ,   .    ,       .   2026    .        -,     ,  .

   ,  ,     ,            .        ,    .   .

   ?

    ,   ,   ,    ,   .           .  ,        ,         .       ,         ,     .

       ,  ,         .     .            ,  ,  ,         .

     ?

,     (, GPT-4),    .     ,        .   ,       . GPT-4   .

  20252026     .

-,      (713  ),    .     ,   80%     ,  ,       .

-,   . ,       ,  10 .      JPEG  500         .        :     16   45 ,      ,     .   ,    16  ,    46          .

   ?

   -,      ,  ,    .      ,   .

   .     ( 48    GGUF)      .    Python        ,      ,     GPU.        .     -    .

     :   .   , ,       .   .

  :     

     .    ,   .      ,    (, , ,       ).              ( 32 000  128 000 )       .           .    .

   ChatGPT,  ,   .            .       ,     .

    

          -.        ,    ,             .        :    .            Python  ,       .

      ,   ,  2026 ,       .




 1.2.  2026      .


              . ,      , GPU      .         ,      -.  ,                 .

   

          API: OpenAI, Anthropic, Google.      ,       .   20232025       : Llama 2, Llama 3, Llama 4  Meta, Mistral  Mixtral   Mistral AI, Qwen 2.5  Alibaba, DeepSeek  .     ,        ,      .

    . Llama 3 8B,   2024 ,    GPT-3.5   ,   2026    713        ,  ,      .       ,   .

     

     : Llama 3 8B   16-    16        .      ,   .

           .  20232024     GGUF  AWQ,      45     . Llama 3 8B    Q4_K_M    5       6        16  .      4060      ,  ,   .

 2026     -.          ,       .       2.2.

   

   .  ,    ,      .         C++,  Python-,    .

 2026    .  ,      :

 llama.cpp    C++         GPU,   .     Raspberry Pi,      ,    .

 llama-cpp-python   Python-  llama.cpp,       Python-,     .

 Ollama  ,     ,      .     .

      llama-cpp-python  ,          Python-       .

  

   .  2026     :

 32     ,         .

   8+  VRAM (NVIDIA RTX 4060/5060, Apple M4  unified memory)     .

         78B     (1020 /).

 ,          .    ,     Python.

    

        .         :

  .       ChatGPT,     .        ,    ,         .

 . GDPR  ,   ,            .    Ȼ   ,    .

   .     ,     ,     .    ,   .

      ,       .     .

    .

              :

    (ChromaDB, LanceDB, Qdrant)        .

   RAG (LangChain, LlamaIndex)       .

    GUI (Tkinter, PySide)        .

   (PyInstaller, Nuitka)  Python-       .

      ,        .

.

2026       :     ,     ,   ,        .       ,    ,   .    .




 2.  :    





 2.1.    2026 : Llama, Mistral, Qwen, Gemma


     ,     :     ?     ,   ,    .  2026     ,      .   .

   ?

    ,  :

    (   GGUF),

        ,

       .

      (open-weight models).      ,        .   (GPT-4, Claude, Gemini)          .

    2026 

    ,      .      713          .

Llama (Meta)

  :    Meta (     ),   Llama 1  2023 .  2026   Llama 3.1, Llama 3.2  Llama 4.       ,      .

 :

      .

   ,    (   3.1+).

     (Python, JavaScript, TypeScript).

  :  ,         Llama.

         (, ,  ).

 :

        Qwen.

       .

   : Llama 3.1 8B ( )  Llama 4 8B,       .

   ( Q4_K_M):

 :  6 ,  8+ .

 VRAM: 6  (  RTX 3060/4060).

Mistral  Mixtral (Mistral AI)

  :   Mistral AI   2023   Mistral 7B,       .  2024    Mixtral    (mixture of experts),   45B     12B,        .

 :

 : Mistral 7B          /.

 Mixtral 8x7B (   )     70B,    12B.

     (, , ).

 :

       ,   Llama  Qwen (    ).

  Mixtral     ,   .

 : Mistral 7B v0.3      ; Mixtral 8x7B        12+  /VRAM.

  :

 Mistral 7B (Q4_K_M):  6 , VRAM 4+ .

 Mixtral 8x7B (Q4_K_M):  12 , VRAM 8+ .

Qwen (Alibaba)

  :  Qwen   Alibaba   20242025 ,   ,        . Qwen 2.5       .

 :

          2026  (,     Llama 3.1).

    , , .

   ,      .

    : 1.8B, 4B, 7B, 14B, 32B, 72B.

 :

     :    .

       (code-switching),    .

 : Qwen 2.5 7B (   )  Qwen 2.5 14B (  ).

   (Qwen 2.5 7B, Q4_K_M):

 : 6 ,  8+ .

 VRAM: 6 .

Gemma (Google)

  :   Google,    Gemini. Gemma 2 (2024)      ,  ,   .

 :

 : Gemma 2 9B       13B.

     Google (JAX, TensorFlow),      llama.cpp   .

  ,  .

 :

    ,   Qwen  Llama.

        (     Hugging Face!).

      .

 : Gemma 2 9B         .

   (Gemma 2 9B, Q4_K_M):

 : 8 .

 VRAM: 6+ .




 


Llama 3.1 8B      . 8  ,      .    ,   .  68     Q4.           .

Llama 4 8B    Llama 3.1   .   8    68  ,    ,    .   .    Llama 3.1  Llama 4   Llama 4.

Mistral 7B       Mistral AI. 7  ,   6  .     ,    .        , , .

Mixtral 8x7B      .   47  ,     12 .       70B,     12  .     .        .

Qwen 2.5 7B         7  .   Alibaba,      .  68  . ,          .

Qwen 2.5 14B    Qwen  14  .          2026 .  12    .   ,        .

Gemma 2 9B    Google  9  .   ,   8  .     ,   .   ,    ,     .

 ?  

 1:  , 816  ,   

 : Llama 3.1 8B (Q4_K_M)  Qwen 2.5 7B (Q4_K_M).         (1020 /)    .       Qwen.          Llama.

 2:       8+  VRAM

 : Llama 3.1 8B (Q5_K_M)  Qwen 2.5 14B (Q4_K_M).  GPU     4060 /.      (Q5)    (14B).    Qwen,   Llama.

 3:  , 32  , 12+  VRAM

 : Llama 4 8B (Q6_K)  Qwen 2.5 14B (Q5_K_M)  Mixtral 8x7B.           .       .

    ?

     Llama 3.1 8B ( Q4_K_M). :

 :       .

  : llama.cpp, Ollama, LangChain     Llama   .

  :      ,       .

       ,   ,    ,     ,   Llama  Qwen         .

 : Llama 3.1 8B Instruct,  GGUF,  Q4_K_M,   ~5 .        .

,   ,                .




 2.2.       


      Llama 3.1 8B      .       ,     ,       Q4_K_M, Q5_K_S, Q8_0  .    ?   ?      16    5     ?  .

       

            ,     .       FP16 (16-    ).     2 .  8    2    16 ,    Llama 3 8B  .

 :  FP16       45   .    ,     . ,              ,    35    .

  :   

,         24  / 192 .    .      MP3   320 /.    ,     10 .

      .   16-      4-, 5-  8-.      ,      .            .

 : GGUF  AWQ

 2026       :

 GGUF (GPT-Generated Unified Format)   GGML,          llama.cpp.      :  ,         Python-.

 AWQ (Activation-aware Weight Quantization)    ,  GPU  .        ,   .     ,    .

 ,    ,   .gguf.

 :   Q4_K_M

    Llama 3.1 8B  Hugging Face.   - :

text

llama-3.1-8b-instruct-Q2_K.gguf

llama-3.1-8b-instruct-Q3_K_S.gguf

llama-3.1-8b-instruct-Q4_K_M.gguf

llama-3.1-8b-instruct-Q5_K_M.gguf

llama-3.1-8b-instruct-Q8_0.gguf

      Q4_K_M:

 Q4    .       4 .     :   ,   ,     .

 K      K-quant,        .     ,    .         .

 M   (Medium).  S (Small), M (Medium), L (Large).   ,     :

o S      .  ,    .

o M    (   ).

o L   ,  ,   .

 , Q4_K_M   4-    K-quant    .     .




     


Q2_K   .  Llama 3.1 8B   3.5    4  .   :   ,    .      , ,  Raspberry Pi    .

Q3_K_S   .   4 ,  5  .     : ,  ,  .     ,  Q4  .

Q4_K_S      .  4.5 ,  5.5  .  ,    Q4_K_M    . ,           .

Q4_K_M   .    5 ,  6  .   ,         .         ,   .

Q5_K_M   .  6 ,  7  .  ,       . ,             8B-.

Q6_K   .  7 ,  8  .            .   GPU  8  VRAM,           .

Q8_0    .  9 ,  10  .      16- .    GPU  12    VRAM,     .

 :   ,     Q4_K_M.   ,         .

    :    

   , ,     .   .

  Q4_K_M    16-      ,   ,    .               ,   95%      .

  Q2_K   :   ,    ,    .     ,     .

  Q5_K_M    :       .        ,   320 /    lossless   .

 :      

1.      .     VRAM     ?  ,   ,        .    ,     .

2.     .     ,            12 .

3.      K-quant.    K (Q4_K_M, Q5_K_M)       K (, Q4_0),        .

4.      Q4_K_M.    ,   ,    Q4_K_M.   ,     .                   .

     

     Llama 3.1 8B Instruct,  GGUF,  Q4_K_M.     5 .   :

   ,   -   ,

     8    6  VRAM,

   1530 /     60 /  ,

     RAG   .

   ,              ,    .

,   ,     ,   .      ,     .

 2.3. :    

  ,      ,    : llama-3.1-8b-instruct-Q4_K_M.gguf.       . ,   Hugging Face  ,   ,     ,    5 .    ,    :  Python-,     .

  ?      -        .    ,  ,   ,     .     , ...   .      .

 

  Python (    ,  3.10  )   huggingface_hub.     Hugging Face      .  :

bash

pip install huggingface_hub

 

       download_model().     :

1.       (, "unsloth/Llama-3.1-8B-Instruct")    ("Llama-3.1-8B-Instruct-Q4_K_M.gguf").

2.    ,      .     .

3.         ,    .

4.        ,      .

     model_loader.py    :

python# model_loader.pyimport osimport sysfrom pathlib import Pathfrom huggingface_hub import hf_hub_downloaddef download_model(    repo_id: str = "unsloth/Llama-3.1-8B-Instruct",    filename: str = "Llama-3.1-8B-Instruct-Q4_K_M.gguf",    model_dir: str = "./models") -> str:    """       GGUF  Hugging Face,     .     :        repo_id:    Hugging Face (, "unsloth/Llama-3.1-8B-Instruct")        filename:             model_dir:          :                  """    #  ,       Path(model_dir).mkdir(parents=True, exist_ok=True)     #  ,        local_path = Path(model_dir) / filename     #            if local_path.exists():        print(f"   : {local_path}")        return str(local_path.resolve())     #       print(f"   {filename}  {repo_id}...")    print("    ~5 .      .")     try:        #           downloaded_path = hf_hub_download(            repo_id=repo_id,            filename=filename,            local_dir=model_dir,            resume_download=True,  #             )        print(f"   : {downloaded_path}")        return downloaded_path     except KeyboardInterrupt:        print("\n   .         .")        sys.exit(1)    except Exception as e:        print(f"   : {e}")        print("   -   repo_id/filename.")        sys.exit(1)if __name__ == "__main__":    #  :    ,        path = download_model()    print(f"\n!   : {path}")

 

 from huggingface_hub import hf_hub_download   ,     :    ,  ,   ,    .

  if local_path.exists()   .   5 ,          .   ,     .      .

 resume_download=True     .   ,        ,     . huggingface_hub        .

  KeyboardInterrupt    ,   Ctrl+C.   ,          .



      :

bashpython model_loader.py    :text   Llama-3.1-8B-Instruct-Q4_K_M.gguf  unsloth/Llama-3.1-8B-Instruct...    ~5 .      .  Downloading: 100%|| 5.12G/5.12G [05:30<00:00, 15.5MB/s]   : models/Llama-3.1-8B-Instruct-Q4_K_M.gguf!   : models/Llama-3.1-8B-Instruct-Q4_K_M.gguf  :text   : models/Llama-3.1-8B-Instruct-Q4_K_M.gguf!   : models/Llama-3.1-8B-Instruct-Q4_K_M.gguf

  

  models/      .   ,     .     ,    ,      -  .




    


 Repo not found. ,      .     huggingface.co (https://huggingface.co/)  ,        .       : unsloth, bartowski, TheBloke, mradermacher.     GGUF-,      .

 Entry not found. ,       .      ,     .                  .

  .    -       Hugging Face.         :    HF_ENDPOINT=https://hf-mirror.com.     ,     .

    .   Llama 3.1 8B   5 ,     hf_hub_download   .     6   .         .

   .          , ,     .     ,     .       ,        .

    ,    Ollama

    Ollama   ,     ollama pull llama3.1:8b.      ?

  ,  Ollama    ,           HTTP.    ,          EXE- (        Ollama).     :     ,   Python-  .           .

            ,         .



   huggingface_hub         .

   download_model(),         .

           .

      models/    .

    .   3    .




 3. :     





 3.1.  : Ollama, llama.cpp, llama-cpp-python


         .    : ,      ,   ,       .          .     ,             .   .

    LLM

 (runtime)     .    ,     GGUF   :       ,  ,      .        .     .

     LLM       llama.cpp     C++,      .  ,       ,   Python, .

 1:  llama.cpp (C++)

  : ,   C++,  - .   ,           :

bash

./llama-cli -m models/Llama-3.1-8B-Instruct-Q4_K_M.gguf -p ", !"

   Python:      (subprocess).     stdin     stdout.

:

   ( C++  ).

      (  , , ).

    (     ).

:

     .

   API  Python:    ,   .

    :      .

     (streaming).

:       ,            Python-.

 2: Ollama

Ollama    ,      (   Windows, macOS, Linux).    ,       REST API,   OpenAI.

  Ollama  :

bash

ollama pull llama3.1:8b

ollama run llama3.1:8b

  Python      HTTP:

python

import requests

response = requests.post("http://localhost:11434/api/generate", json={

"model": "llama3.1:8b",

"prompt": ", !"

})

:

         beginner-friendly.

    (      2.3).

   OpenAI API:    openai  Python,   base_url  localhost.

  CLI  - ( ).

:

   ,    .      .

     EXE    .     Ollama,   .

  : Ollama   , ,  .       .

   (,       )    API.

:       .        ,  Ollama.   ,         ,    -  .

 3: llama-cpp-python ( )

 Python-  llama.cpp.  , llama.cpp    ,  llama-cpp-python  Python-    .  :

bash

pip install llama-cpp-python

   Python:

python

from llama_cpp import Llama

model = Llama(model_path="models/Llama-3.1-8B-Instruct-Q4_K_M.gguf")

response = model.create_chat_completion(

messages=[{"role": "user", "content": ", !"}]

)

print(response["choices"][0]["message"]["content"])

:

    .      Python-.   , HTTP-,  .

  .    llama.cpp: , top_p,  ,  ,  (CPU/CUDA/Metal/Vulkan).     .

   OpenAI API. create_chat_completion     ,    .       openai.OpenAI(api_key=...)   Llama(model_path=...).

  .   streaming  generator.

   EXE. PyInstaller   llama-cpp-python    llama.cpp    .     .

  .      llama.cpp,      .

:

    GPU     (   3.2).

    ,    llama.cpp (   510%).

            (      ).

:     . lloma-cpp-python        Python-,           .




    


llama.cpp ()

      C++       .   Python       ,    .    ,     ,     .   OpenAI API .    EXE      .  GPU   .     ,    ,    .

Ollama

           .  HTTP API,        .        ,       EXE-.       Ollama     .   ,   OpenAI API . GPU   .     ,   .

llama-cpp-python

   pip install.   Python-  llama.cpp,      ,        .     ,       OpenAI API   create_chat_completion.    EXE  PyInstaller   .  GPU    .         ,      .

    

    llama-cpp-python.   :

1.     .      Python.    Ollama  .   .

2.     .      EXE,       .      .  Ollama  ,   llama.cpp   .

3.    .      ,  ,   (CPU  GPU)    ,     . Ollama       ,    .

4.     API. create_chat_completion     ,    OpenAI.   -    ChatGPT,   ,      .

 -   Ollama

    ,  Ollama   .     :

            .

     (Open WebUI, AnythingLLM),   Ollama  .

     API   .

   ,   EXE -   llama-cpp-python.

 

     llama-cpp-python,      GPU,    ,    .   - LocalModel,      .




 3.2.    llama-cpp-python


  .     ,       .  llama-cpp-python     pip install.     (     ),    .            ,   .

     

llama-cpp-python   Python-.       llama.cpp.    pip        .     ,            .

    :

  CPU.  ,  .       .

   GPU.    SDK   .    ,    .

 :    CPU

  ,             .    Windows, macOS  Linux.

bash

pip install llama-cpp-python

.          CPU.     .    macOS  Apple Silicon (M1/M2/M3/M4),       Accelerate (Apple Neural Engine),     .

   GPU:   

     ,      .      310 .

  NVIDIA (CUDA)

 :

  NVIDIA ( ,    nvidia.com (https://nvidia.com/)).

 CUDA Toolkit 12.x (  developer.nvidia.com/cuda-downloads (https://developer.nvidia.com/cuda-downloads)).

 CUDA Toolkit,  :

bash

CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python

 CMAKE_ARGS   -DGGML_CUDA=on .    CUDA  llama.cpp.    .

:    Python  :

python

from llama_cpp import Llama

#      "CUDA not available",  .

    GGML_CUDA=off, ,   .   (pip uninstall llama-cpp-python)   .

 macOS  Apple Silicon (Metal)

  Mac   M1/M2/M3/M4   GPU-  Metal.    CPU-   Accelerate,     Metal   :

bash

CMAKE_ARGS="-DGGML_METAL=on" pip install llama-cpp-python

       Apple,     CPU.

  AMD (Vulkan)

 AMD    Vulkan:

bash

CMAKE_ARGS="-DGGML_VULKAN=on" pip install llama-cpp-python

  ,     Vulkan SDK (vulkan.lunarg.com (https://vulkan.lunarg.com/)).




     .


  (Intel/AMD).     .    pip install llama-cpp-python   .      ,          515     8B-. ,             .

NVIDIA GeForce/RTX.     NVIDIA   CUDA,      .  : CMAKE_ARGS="-DGGML_CUDA=on" pip install llama-cpp-python.   ,    NVIDIA  CUDA Toolkit.     3060   ,          .

Apple M1/M2/M3/M4.   Mac   Apple Silicon    GPU-  Metal.       pip install llama-cpp-python     Accelerate.       -DGGML_METAL=on.      ,        unified memory.

AMD Radeon.  AMD    Vulkan.  : CMAKE_ARGS="-DGGML_VULKAN=on" pip install llama-cpp-python.    Vulkan SDK   vulkan.lunarg.com (https://vulkan.lunarg.com/).    ,        .

:    ,        pip install llama-cpp-python  .     ,   .        GPU.

 :  Hello, world  

    test_model.py      :

python

# test_model.py

from model_loader import download_model

from llama_cpp import Llama

# 1.  ()

model_path = download_model()

# 2. 

print("   ...")

model = Llama(

model_path=model_path,

n_ctx=2048,        #    (   )

n_threads=4,       #   CPU (   )

verbose=False      #   

)

# 3.  

print(" ...")

response = model.create_chat_completion(

messages=[

{"role": "system", "content": "   .  ."},

{"role": "user", "content": "     ?"}

],

temperature=0.7,

max_tokens=100

)

# 4. 

answer = response["choices"][0]["message"]["content"]

print(f"\n :\n{answer}")

:

bash

python test_model.py

   ,   - :

text

  Llama-3.1-8B-Instruct-Q4_K_M.gguf  unsloth/Llama-3.1-8B-Instruct...

  : models/Llama-3.1-8B-Instruct-Q4_K_M.gguf

   ...

 ...

 :

     ,               .

    Llama()

 model_path     .       download_model().

 n_ctx      .    ,   .   2048   (  1500 ). ,   RAG,   4096  8192.

 n_threads    CPU,    .      .  4-   4,  8-  8.

 verbose=False     .  -   ,    (True)  .




    


ModuleNotFoundError: No module named 'llama_cpp'.       . ,    venv   pip install llama-cpp-python.   ,     ,      :        (venv).

ImportError: DLL load failed.    Windows    Visual C++.    Microsoft Visual C++ Redistributable    Microsoft.        .

RuntimeError: CUDA error: no CUDA-capable device is detected.      CUDA,   NVIDIA  .      NVIDIA,    . :     nvidia.com (https://nvidia.com/),   llama-cpp-python   CUDA    .

Illegal instruction (core dumped).   Linux,      AVX2,  llama.cpp   .       .   -DGGML_NATIVE=off         ,    .

    12   .  ,       .   n_threads    Llama          .     ,     GPU     .

  out of memory.        .    :     (, Q3_K_M  Q4_K_M    ),      (7B  8B).    ,    , IDE, .

    

  llama-cpp-python   ,     .    ,      GPU,     CPU.  ,       :

python

model = Llama(

model_path=model_path,

n_ctx=2048,

n_threads=4,

n_gpu_layers=20  #  20    GPU,    CPU

)

 n_gpu_layers    0 (  CPU)  -1 (  GPU,   ).    ,          .

  

    :

  model_loader.py    .

  test_model.py         .

    :   - LocalModel,        (,  , RAG, ).      3.3.




 3.3.  - LocalModel   chat()  stream().


  llama-cpp-python, ,        .        ,  .    LocalModel,     .     ,   ,       .

  

  Llama()  llama-cpp-python ,     :

       (  ,  , ).

     .

 ,    ,   - .

      (      ).

  LocalModel   .  :

      ,

    chat()  stream_chat(),

     (),

     ,   OpenAI API,     ,    ChatGPT.

  LocalModel

    local_model.py.        .

python

# local_model.py

from typing import List, Dict, Optional, Generator

from llama_cpp import Llama

class LocalModel:

"""

  llama-cpp-python,   

     .

"""

def __init__(

self,

model_path: str,

n_ctx: int = 2048,

n_threads: int = 4,

n_gpu_layers: int = 0,

verbose: bool = False

):

"""

 .

:param model_path:    .gguf

:param n_ctx:    (  ,

  )

:param n_threads:   

:param n_gpu_layers:     GPU (-1 = , 0 = )

:param verbose:     

"""

self.model_path = model_path

self.n_ctx = n_ctx

#    

self._llm = Llama(

model_path=model_path,

n_ctx=n_ctx,

n_threads=n_threads,

n_gpu_layers=n_gpu_layers,

verbose=verbose

)

#    (  id )

self._histories: Dict[str, List[Dict[str, str]]] = {}

def chat(

self,

messages: List[Dict[str, str]],

temperature: float = 0.7,

max_tokens: int = 512,

system_prompt: Optional[str] = None

) -> str:

"""

      .

:param messages:     OpenAI:

[{"role": "user", "content": "..."}, ...]

:param temperature:  (0  ,  1.5  )

:param max_tokens:   

:param system_prompt:   (    messages)

:return:    

"""

#   system_prompt,     

full_messages = []

if system_prompt:

full_messages.append({"role": "system", "content": system_prompt})

full_messages.extend(messages)

response = self._llm.create_chat_completion(

messages=full_messages,

temperature=temperature,

max_tokens=max_tokens

)

return response["choices"][0]["message"]["content"]

def stream_chat(

self,

messages: List[Dict[str, str]],

temperature: float = 0.7,

max_tokens: int = 512,

system_prompt: Optional[str] = None

) -> Generator[str, None, None]:

"""

  chat().  ,  

    .   

 ,   ChatGPT.

:param messages:  

:param temperature: 

:param max_tokens:   

:param system_prompt:  

:yield:   ()

"""

full_messages = []

if system_prompt:

full_messages.append({"role": "system", "content": system_prompt})

full_messages.extend(messages)

stream = self._llm.create_chat_completion(

messages=full_messages,

temperature=temperature,

max_tokens=max_tokens,

stream=True

)

for chunk in stream:

#         

if "choices" in chunk and len(chunk["choices"]) > 0:

delta = chunk["choices"][0].get("delta", {})

content = delta.get("content", "")

if content:

yield content

def save_history(self, conversation_id: str, messages: List[Dict[str, str]]):

"""     id."""

self._histories[conversation_id] = messages.copy()

def load_history(self, conversation_id: str) -> List[Dict[str, str]]:

"""    id.     ."""

return self._histories.get(conversation_id, [])

def clear_history(self, conversation_id: str):

"""  ."""

if conversation_id in self._histories:

del self._histories[conversation_id]

  

 __init__

       . ,         LocalModel       .      .  n_ctx   :   ,       ,     .   2048  . n_gpu_layers      .

 chat()

     OpenAI   ,    API.  ,      ,   ChatGPT,      LocalModel,    .     ,   ,   create_chat_completion.   ,   .

 stream_chat()

         ,   ChatGPT.      ,       .  Python     (  yield).     :

python

for token in model.stream_chat(messages):

print(token, end="", flush=True)

         .        GUI.

 

   _histories,       .  ,        .        (   ),    .

 :  

  chat_demo.py:

python

# chat_demo.py

from model_loader import download_model

from local_model import LocalModel

# 1.    

model_path = download_model()

print(" ...")

model = LocalModel(model_path=model_path, n_threads=4)

# 2.   (,  )

system = "  -.  ,   ."

# 3.   

conversation_id = "cooking"

messages = []  #    

while True:

user_input = input("\n: ")

if user_input.lower() in ["", "quit", "exit"]:

break

#     

messages.append({"role": "user", "content": user_input})

#   ()

print(": ", end="", flush=True)

full_response = ""

for token in model.stream_chat(

messages=messages,

system_prompt=system,

temperature=0.7

):

print(token, end="", flush=True)

full_response += token

print()  #    

#    

messages.append({"role": "assistant", "content": full_response})

model.save_history(conversation_id, messages)

:

bash

python chat_demo.py

 :

text

:   ?

:   4    .   8-10 .    ,   .

:  ?

:    ,  ,  2-3 .      ,    .

 :  ,    

  n_ctx    ,     .     ,   .       N .    trim_history    (  local_model.py):

python

def trim_history(self, messages: List[Dict[str, str]], max_messages: int = 20) -> List[Dict[str, str]]:

"""

   max_messages .

  (   role='system')  .

"""

if len(messages) <= max_messages:

return messages

#   ,   

system_messages = [m for m in messages if m["role"] == "system"]

other_messages = [m for m in messages if m["role"] != "system"]

#  

trimmed = other_messages[-max_messages:]

return system_messages + trimmed




    .


 ,    .  ,       n_ctx. Llama.cpp   ,    .  :  n_ctx     ,  2048  4096 . ,       .   n_ctx  -  ,   trim_history()       N ,   .

    .       max_tokens         .  max_tokens  500  1000.      . ,      flush=True: print(token, end="", flush=True).   Python   ,     .

     .           ,  . ,       ,  :      .     ,   .      RAG   4 :       ,   .

 KeyError: 'choices'.     ,     . Llama.cpp  chunk   .         : if "choices" in chunk and len(chunk["choices"]) > 0:.  ,        .

 .

    ,       LLM.              Python-.      ,          HTTP-,   OpenAI API.          ,     ChatGPT.




 3.4.  API   OpenAI


   LocalModel,        Python-.  ,         -  ? ,   VS Code, -,      ,   ChatGPT?        .

 ,     . OpenAI  HTTP API,        .         API.     ,   OpenAI,        .        .

  

   HTTP-  Python, :

   (, 8080)   .

  POST-   /v1/chat/completions    ,   OpenAI.

      LocalModel.

      JSON-,   OpenAI.

    (streaming),     .

      API     .

    , ,   Python- openai,   base_url="http://localhost:8080/v1".         Llama,   ChatGPT. !



  HTTP-    FastAPI.  ,    ,     API.    ,     .    uvicorn  ,     FastAPI-.

 :

bash

pip install fastapi uvicorn

 

  openai_api.py   .       .

python

# openai_api.py

import time

import uuid

from typing import Optional, List, Dict, Any

from fastapi import FastAPI, HTTPException

from fastapi.responses import StreamingResponse

from pydantic import BaseModel, Field

import uvicorn

#    LocalModel    

from local_model import LocalModel

from model_loader import download_model

# ---  ,   OpenAI API ---

class ChatMessage(BaseModel):

"""  ."""

role: str  # "system", "user", "assistant"

content: str

class ChatCompletionRequest(BaseModel):

""" ,   OpenAI."""

model: str = "local-model"

messages: List[ChatMessage]

temperature: Optional[float] = 0.7

max_tokens: Optional[int] = 512

stream: Optional[bool] = False

class ChatCompletionChoice(BaseModel):

"""  ."""

index: int

message: ChatMessage

finish_reason: Optional[str] = "stop"

class ChatCompletionResponse(BaseModel):

"""  ( )."""

id: str

object: str = "chat.completion"

created: int

model: str

choices: List[ChatCompletionChoice]

# ---  FastAPI   ---

app = FastAPI(title="Local AI API", description="OpenAI- API   LLM")

#       

model_path = download_model()

print("   ...")

local_model = LocalModel(model_path=model_path, n_threads=4, n_gpu_layers=0)

print("   .")

# ---   ---

def messages_to_list(messages: List[ChatMessage]) -> List[Dict[str, str]]:

""" Pydantic-     LocalModel."""

return [{"role": m.role, "content": m.content} for m in messages]

def generate_stream_response(messages: List[Dict[str, str]], temperature: float, max_tokens: int):

"""      OpenAI (Server-Sent Events)."""

chat_id = f"chatcmpl-{uuid.uuid4().hex[:12]}"

created = int(time.time())

#     

for token in local_model.stream_chat(

messages=messages,

temperature=temperature,

max_tokens=max_tokens

):

chunk = {

"id": chat_id,

"object": "chat.completion.chunk",

"created": created,

"model": "local-model",

"choices": [

{

"index": 0,

"delta": {"content": token},

"finish_reason": None

}

]

}

yield f"data: {__import__('json').dumps(chunk)}\n\n"

#    finish_reason

final_chunk = {

"id": chat_id,

"object": "chat.completion.chunk",

"created": created,

"model": "local-model",

"choices": [

{

"index": 0,

"delta": {},

"finish_reason": "stop"

}

]

}

yield f"data: {__import__('json').dumps(final_chunk)}\n\n"

yield "data: [DONE]\n\n"

# ---  endpoint ---

@app.post("/v1/chat/completions", response_model=None)

async def chat_completions(request: ChatCompletionRequest):

"""

 endpoint,   OpenAI Chat Completions API.

  ,    .

"""

#  

messages = messages_to_list(request.messages)

#    

if request.stream:

return StreamingResponse(

generate_stream_response(

messages=messages,

temperature=request.temperature or 0.7,

max_tokens=request.max_tokens or 512

),

media_type="text/event-stream"

)

#  

try:

response_text = local_model.chat(

messages=messages,

temperature=request.temperature or 0.7,

max_tokens=request.max_tokens or 512

)

except Exception as e:

raise HTTPException(status_code=500, detail=f" : {str(e)}")

#     OpenAI

response = ChatCompletionResponse(

id=f"chatcmpl-{uuid.uuid4().hex[:12]}",

created=int(time.time()),

model=request.model,

choices=[

ChatCompletionChoice(

index=0,

message=ChatMessage(role="assistant", content=response_text),

finish_reason="stop"

)

]

)

return response

@app.get("/v1/models")

async def list_models():

"""    ( )."""

return {

"object": "list",

"data": [

{

"id": "local-model",

"object": "model",

"created": int(time.time()),

"owned_by": "local"

}

]

}

if __name__ == "__main__":

#     8080

uvicorn.run(app, host="127.0.0.1", port=8080)

  

  (Pydantic)

  pydantic.BaseModel      .  ,   API     OpenAI.  ,      ChatGPT,     :  model, messages, temperature, max_tokens, stream  .

 

   "stream": true,      JSON.     Server-Sent Events (SSE)   ,   data:.      JSON-   .     data: [DONE],   .

 generate_stream_response   ,  yield . FastAPI    StreamingResponse   media type,      ,   ChatGPT.

   

 : LocalModel   ,   .         .  ,     ,    .



 :

bash

python openai_api.py

 :

text

   ...

   .

INFO:     Started server process [12345]

INFO:     Waiting for application startup.

INFO:     Application startup complete.

INFO:     Uvicorn running on http://127.0.0.1:8080 (Press CTRL+C to quit)

  API   curl  Python.  - :

python

# test_openai_api.py

import requests

response = requests.post(

"http://127.0.0.1:8080/v1/chat/completions",

json={

"model": "local-model",

"messages": [

{"role": "user", "content": "!  ?"}

],

"temperature": 0.7,

"max_tokens": 100,

"stream": False

}

)

print(response.json()["choices"][0]["message"]["content"])

   :

python

import requests

response = requests.post(

"http://127.0.0.1:8080/v1/chat/completions",

json={

"model": "local-model",

"messages": [

{"role": "user", "content": " "}

],

"stream": True

},

stream=True

)

for line in response.iter_lines():

if line:

line = line.decode("utf-8")

if line.startswith("data: ") and not line.endswith("[DONE]"):

import json

chunk = json.loads(line[6:])

token = chunk["choices"][0]["delta"].get("content", "")

if token:

print(token, end="", flush=True)

print()

    OpenAI

  .     openai,      ChatGPT,     URL    (  ):

bash

pip install openai

python

# test_with_openai_lib.py

from openai import OpenAI

client = OpenAI(

base_url="http://127.0.0.1:8080/v1",

api_key="not-needed"  #    ,     

)

# - 

response = client.chat.completions.create(

model="local-model",

messages=[{"role": "user", "content": "  Python?"}],

temperature=0.7,

max_tokens=100

)

print(response.choices[0].message.content)

#  

stream = client.chat.completions.create(

model="local-model",

messages=[{"role": "user", "content": "   "}],

stream=True

)

for chunk in stream:

if chunk.choices[0].delta.content:

print(chunk.choices[0].delta.content, end="", flush=True)

print()

.

    127.0.0.1 (localhost).  ,          .        ,  host="127.0.0.1"  host="0.0.0.0",     (,  API-).          .




    API.


 Address already in use.  8080,     ,    .         : port=8081.     8080,  ,  ,   .  Windows    netstat -ano | findstr :8080,  taskkill /PID _ /F.

    .          ,        .          :  LocalModel   ,    -.        ,    .

     .    ,     .  max_tokens,      .    temperature  0.30.5            .           .

     .  API       OpenAI:   id, object, created, model, choices .  -       ,  (,  openai)   .              OpenAI.

  .

     HTTP API, :

    .

   OpenAI Chat Completions API.

   .

      ,  OpenAI API (  ).

        .

  .       ,         .        (RAG, ),  API     .

  4             .




 4.   :    





 4.1.   RAG  


         , ,  .    ,  ,   ChatGPT:    ,   .   :     ""    15 ?,    ,   .     ,    PDF-,      .

       ,   ,      .    RAG  Retrieval-Augmented Generation,  ,  .      ,      .

:     

       ,   .             .         ,   ,     :        .

RAG   .        ,              .

:    

,         ,   .       ,       .    ,    -   ,      :       .           .

  RAG:

    ,   .

      (),     .

  ,         ( ).

  ,        (LLM).

      .       ,   .

  :  

   .         .

 1:  ( )

       PDF, Word,  , Excel-        .

      (chunks)     5001000 .   ?       ,    200-    .       .

     embedding-   ,      ( ).    .         ,      .

       .   ,      ,   .    ChromaDB  ,     .

     (   )       .

 2:  ()

   , :        ""?,  :

       embedding-.

        ,      .   .

 , , 4   :     ,           .

 3:   (augmentation)

           .    :

text

  ,        .

  ,    .   ,  "   ".

:        ""?

:

---  1 (   "".pdf,  7):

7.1.       30 ...

---  2 (  .docx,  3):

...

:

        .        .

 4:  

  ,     .    :   7.1    ""...            .

    

  RAG     :

1.    Embedding-.    sentence-transformers    intfloat/multilingual-e5-base  BAAI/bge-base-en.    ~250500      CPU.

2.      . ChromaDB      ,         Python.

3.     .  Llama 3.1,   LocalModel.

4.    .              Python-,   .

        .  embedding-,     ,  .

RAG     ChatGPT

 -   PDF   .  , :

 :     .        .

 :           .

 :    .

  RAG   .           .

 RAG,    

RAG    .     :

  .  embedding-       ,     .    ,        .

  .      (      ).           .      4.4.

 .       .     ,     .

     

    4       RAG:

  4.2  embedding-    .

  4.3   ChromaDB   ,  ,  .

  4.4       (PDF, DOCX, TXT).

  4.5   ask_documents()     ,   ,       .

         .          ,    ,      .




 4.2.   embedding- (BGE, multilingual-e5, jina)


    ,  RAG          ,  .    embedding-.   ,      ,      .     ,      ,   ,      .

  embedding-

Embedding-     (     )        ,   384  1024 .  : ,   ,  ,          .       .

:

             .

              .

        : BGE, multilingual-e5  jina.     Hugging Face,    sentence-transformers      .

    

     :

1.   .       .      ,      .

2.   .   ,      .    .

3.   .     CPU  GPU,    (200500 )    . : embedding-      ( ),     ( ).

4. .       sentence-transformers,        Windows, macOS  Linux.



BGE (BAAI General Embedding)

       (BAAI).  : bge-base-en-v1.5 (), bge-m3 ().

 bge-base-en-v1.5:  ,  438 ,    768.    ,    .     .

 bge-m3:  ,   100 ,  .  ~2.2 ,   1024.    ,        .      CPU       .  ,     :     100 .

:  ,          .

Multilingual-E5 (Microsoft/Intfloat)

 multilingual-e5     Microsoft  . : multilingual-e5-small, multilingual-e5-base, multilingual-e5-large.

 multilingual-e5-small:  ~470 ,    384.   90 ,  .

 multilingual-e5-base:  ~1.1 ,   768.    ,   small.

 multilingual-e5-large:  ~2.2 ,   1024.  ,   .

  (small  base)     ,    .  :   E5       "query: "     "passage: "   .   ,    .

: multilingual-e5-base    :    (1.1 )  . multilingual-e5-small     .

Jina AI Embeddings

 Jina AI   jina-embeddings-v3,     jina-embeddings-v2.   ,   .

 jina-embeddings-v2-base-multilingual:  ~550 ,   768.  , ,     .

 jina-embeddings-v3:     task-specific embeddings (  ),  ~1.5 ,    BGE-m3.

 Jina        ( 8192 ),     .

: jina-embeddings-v2-base-multilingual  , ,    .    E5.




 embedding-  RAG


bge-m3.    BAAI    1024   2.2 .     ,                GPU.    ,        .

multilingual-e5-small.        470     384.    ,       .    Raspberry Pi,    ,     .      .

multilingual-e5-base.       .  1.1 ,   768.        .            ,      .    .

multilingual-e5-large.    multilingual-e5.  2.2 ,  1024.     ,   ,    .   ,       GPU        ,       .

jina-embeddings-v2-base.    Jina AI.  550    768    e5-base    .        8192 ,     512.  ,             .

 :   multilingual-e5-base          .        e5-small.       e5-large  jina.  embedding-             Embedder.

    

  intfloat/multilingual-e5-base. :

           .        .

  1.1         .    ,     LLM.

       sentence-transformers,   GPU,    .

       "query: "  "passage: "    .   ,   .

     (8    ),      multilingual-e5-small       ,       .

  

   sentence-transformers.     :

bash

pip install sentence-transformers

        .   embedder.py:

python

# embedder.py

from sentence_transformers import SentenceTransformer

class Embedder:

"""

  sentence-transformers     .

"""

def __init__(self, model_name: str = "intfloat/multilingual-e5-base"):

print(f" embedding- {model_name}...")

self._model = SentenceTransformer(model_name)

self._dim = self._model.get_sentence_embedding_dimension()

print(f" ,  : {self._dim}")

@property

def dimension(self) -> int:

""" ."""

return self._dim

def embed_query(self, text: str) -> list:

"""    ."""

#  E5     "query: "

return self._model.encode(f"query: {text}", normalize_embeddings=True).tolist()

def embed_documents(self, texts: list[str]) -> list[list]:

"""   ()   ."""

#     "passage: "

prefixed = [f"passage: {t}" for t in texts]

return self._model.encode(prefixed, normalize_embeddings=True).tolist()

if __name__ == "__main__":

emb = Embedder()

# :      

query_vec = emb.embed_query("  ?")

doc_vecs = emb.embed_documents([

"     12 .",

"       .",

"       0.1%   ."

])

print(f" , : {len(query_vec)}")

print(f" : {len(doc_vecs)}")

:

bash

python embedder.py

     ( 1 )  .  :

text

 embedding- intfloat/multilingual-e5-base...

 ,  : 768

 , : 768

 : 3

       .            ChromaDB.




   embedding-


     .     Embedder()  sentence-transformers     Hugging Face.   ,    . :       ,      models/multilingual-e5-base/      .      .

    . Embedding-     . multilingual-e5-base   1.1 ,    .       4       ,    .   multilingual-e5-small    470 ,  ,         .

     .    .      "query: ",      "passage: ".      ,   ,    .  ,      sentence-transformers         .

   .     CPU       .     device="cuda"   SentenceTransformer,         .      CPU-  ,    .

 

   embedding-,     .   -          .     4.3     : ChromaDB  LanceDB.




 4.3.     : ChromaDB  LanceDB


     .     -  , ,    ,   .         ,      .

    Pinecone, Weaviate, Qdrant.  ,         .     , :

   ,     .

      ( Docker, systemd).

       .

    (, ,   embedding-).

   open-source.

 2026        : ChromaDB  LanceDB.     ,     RAG.

ChromaDB:   Python-first

ChromaDB (https://www.trychroma.com/)    ,         Python.       ,           .

:

python

import chromadb

#   (    ./chroma_data)

client = chromadb.PersistentClient(path="./chroma_data")

#   (   SQL)

collection = client.get_or_create_collection(

name="my_docs",

metadata={"hnsw:space": "cosine"}  #     

)

#      

collection.add(

embeddings=[[0.1, 0.2, ...], [0.3, 0.4, ...]],  # 

documents=["  ", "  "],

metadatas=[{"source": "file1.pdf", "page": 1}, {"source": "file2.docx"}],

ids=["doc1_chunk0", "doc2_chunk0"]

)

#    

results = collection.query(

query_embeddings=[[0.15, 0.25, ...]],

n_results=3

)

 :

 ChromaDB:

  .    , API  .

  .        . PersistentClient  ,     Python.

 .       , ,      .     .

   embedding-. ChromaDB    sentence-transformers (     Embedder  ).

  . ChromaDB  ,    .

:

   SQLite3.   ChromaDB  SQLite   .  ,      .

  .    (  )       HNSW.

LanceDB:    

LanceDB (https://lancedb.github.io/lancedb/)      ,      Lance.         .

:

bash

pip install lancedb

 :

python

import lancedb

import pyarrow as pa

#    ( PersistentClient)

db = lancedb.connect("./lancedb_data")

#    

schema = pa.schema([

("vector", pa.list_(pa.float32(), 768)),

("text", pa.string()),

("source", pa.string()),

])

table = db.create_table("my_docs", schema=schema, mode="overwrite")

#  

table.add([{

"vector": [0.1, 0.2, ...],

"text": " ",

"source": "file1.pdf"

}])

# 

results = table.search([0.15, 0.25, ...]).limit(3).to_list()

 LanceDB:

   . LanceDB       .   ,   ChromaDB,    .

   .       .

  Apache Arrow.     Arrow,     - .

  .     ChromaDB, LanceDB      .

:

   .    ,   ChromaDB.

 API  .     PyArrow,      .

   .    embedding-,    .




 ChromaDB  LanceDB


.      pip install. ChromaDB       . LanceDB   Apache Arrow     pyarrow .

 . ChromaDB       ,  SQLite       .           ,   . LanceDB    Lance,        .

API   . ChromaDB     Python-:    add, query, count.      ,    . LanceDB     PyArrow     ,    .

.  ChromaDB     JSON-             .  LanceDB          ,     .     ChromaDB .

.          ,  . ChromaDB       - SQLite  . LanceDB            .

  . ChromaDB  ,    ,      Stack Overflow. LanceDB    ,    .     ChromaDB  -   :       .

 . ChromaDB            Embedder.          .  LanceDB   .

  .       Windows, macOS  Linux  .

: ChromaDB        . LanceDB         .       .

    

  ChromaDB.  :

1.     API.       PyArrow    . ChromaDB      ,    .

2.     .        ,   ,  ,    ,     .  LanceDB       .

3.     .      (,   ) ChromaDB  ,    LanceDB  .

4.       .     ,       ChromaDB.

          ,   LanceDB  Qdrant          .

:  ChromaDB   

  vector_store.py,       :

python

# vector_store.py

import chromadb

from typing import List, Dict, Optional

class VectorStore:

"""

  ChromaDB      .

"""

def __init__(self, collection_name: str = "my_documents", persist_path: str = "./chroma_data"):

self._client = chromadb.PersistentClient(path=persist_path)

self._collection = self._client.get_or_create_collection(

name=collection_name,

metadata={"hnsw:space": "cosine"}

)

def add_documents(

self,

ids: List[str],

embeddings: List[List[float]],

documents: List[str],

metadatas: Optional[List[Dict]] = None

):

"""  ()  ."""

self._collection.add(

ids=ids,

embeddings=embeddings,

documents=documents,

metadatas=metadatas

)

def search(self, query_embedding: List[float], n_results: int = 5) -> Dict:

"""     ."""

return self._collection.query(

query_embeddings=[query_embedding],

n_results=n_results

)

def count(self) -> int:

"""    ."""

return self._collection.count()

def clear(self):

"""  (  )."""

#    

name = self._collection.name

self._client.delete_collection(name)

self._collection = self._client.get_or_create_collection(

name=name,

metadata={"hnsw:space": "cosine"}

)

,    Embedder   4.2:

python

# test_vector_store.py

from embedder import Embedder

from vector_store import VectorStore

# 1.  embedding-

emb = Embedder()

# 2.  

store = VectorStore()

# 3.    

texts = [

"     12    .",

"       .",

"      0.1%    .",

"      30  ."

]

embeddings = emb.embed_documents(texts)

store.add_documents(

ids=[f"doc_{i}" for i in range(len(texts))],

embeddings=embeddings,

documents=texts,

metadatas=[{"source": "test.txt", "chunk": i} for i in range(len(texts))]

)

print(f"  : {store.count()}")

# 4.   

query = "    ?"

query_vec = emb.embed_query(query)

results = store.search(query_vec, n_results=2)

print("\n :")

for i, (doc, meta, dist) in enumerate(zip(

results["documents"][0],

results["metadatas"][0],

results["distances"][0]

)):

print(f"{i+1}. [distance: {dist:.3f}] {doc}")

print(f"   : {meta['source']}\n")

:

bash

python test_vector_store.py

:

text

 embedding- intfloat/multilingual-e5-base...

 ,  : 768

  : 4

 :

1. [distance: 0.123]       30  .

: test.txt

2. [distance: 0.145]      12    .

: test.txt

,       ,      .  (distance)      ,    .

 

      RAG: embedding-    .      ,     ,   PDF, DOCX, TXT,        ChromaDB.    4.4 :    .




 4.4. :    .


       : embedding-    , ChromaDB    .      ,       ,  ,     ,      .      ,     ( )         .

    

       :

 PDF  , ,  ( ).    PyPDF2.

 DOCX   Microsoft Word.  python-docx.

 TXT     (   Python).

  :

bash

pip install PyPDF2 python-docx langchain-text-splitters

langchain-text-splitters    ,     .      LangChain,   langchain-core.    RecursiveCharacterTextSplitter       ,   ,   ,    .

  :   

    ChromaDB  100- PDF   ,             .     ,      .

     5001000 .     ,            LLM.     (overlap)  100200 ,     .    :      ,     .

  

  document_reader.py:

python

# document_reader.py

import os

from typing import List, Dict

import PyPDF2

from docx import Document

def read_pdf(file_path: str) -> str:

"""   PDF-,   ."""

text_parts = []

with open(file_path, 'rb') as f:

reader = PyPDF2.PdfReader(f)

for page_num, page in enumerate(reader.pages, start=1):

page_text = page.extract_text()

if page_text:

text_parts.append(f"[ {page_num}] {page_text}")

return "\n".join(text_parts)

def read_docx(file_path: str) -> str:

"""   DOCX-."""

doc = Document(file_path)

text_parts = [para.text for para in doc.paragraphs if para.text.strip()]

return "\n".join(text_parts)

def read_txt(file_path: str) -> str:

"""      ."""

#  UTF-8,     cp1251 (Windows-)

try:

with open(file_path, 'r', encoding='utf-8') as f:

return f.read()

except UnicodeDecodeError:

with open(file_path, 'r', encoding='cp1251') as f:

return f.read()

def read_file(file_path: str) -> str:

"""

         .

"""

ext = os.path.splitext(file_path)[1].lower()

if ext == '.pdf':

return read_pdf(file_path)

elif ext == '.docx':

return read_docx(file_path)

elif ext == '.txt':

return read_txt(file_path)

else:

raise ValueError(f" : {ext}")

:

 read_pdf         .

 read_docx    ,      .

 read_txt   :    Windows-1251 (    ).

  

    indexer.py,    :

python

# indexer.py

import os

import uuid

from typing import List, Dict

from langchain_text_splitters import RecursiveCharacterTextSplitter

from document_reader import read_file

from embedder import Embedder

from vector_store import VectorStore

def index_folder(

folder_path: str,

store: VectorStore,

embedder: Embedder,

chunk_size: int = 800,

chunk_overlap: int = 150,

file_extensions: List[str] = ['.pdf', '.docx', '.txt']

) -> int:

"""

  ,    ,

  ,     ChromaDB.

    .

"""

#   

splitter = RecursiveCharacterTextSplitter(

chunk_size=chunk_size,

chunk_overlap=chunk_overlap,

separators=["\n\n", "\n", ". ", " ", ""]  # :     

)

total_chunks = 0

#       

for root, dirs, files in os.walk(folder_path):

for file_name in files:

ext = os.path.splitext(file_name)[1].lower()

if ext not in file_extensions:

continue  #   

file_path = os.path.join(root, file_name)

print(f": {file_path}")

try:

#    

full_text = read_file(file_path)

if not full_text.strip():

print(f"  ->  , .")

continue

#   

chunks = splitter.split_text(full_text)

print(f"  ->  : {len(chunks)}")

#     

ids = []

embeddings_list = []

documents = []

metadatas = []

for i, chunk in enumerate(chunks):

chunk_id = f"{file_path}_{i}_{uuid.uuid4().hex[:8]}"

ids.append(chunk_id)

documents.append(chunk)

metadatas.append({

"source": file_path,

"file_name": file_name,

"chunk_index": i

})

#     (batch-)

embeddings_list = embedder.embed_documents(chunks)

#   ChromaDB

store.add_documents(

ids=ids,

embeddings=embeddings_list,

documents=documents,

metadatas=metadatas

)

total_chunks += len(chunks)

except Exception as e:

print(f"  -> : {e}")

continue

return total_chunks

:

   RecursiveCharacterTextSplitter,         (),   ,     .     .

  (source, file_name, chunk_index)     ,        .

   : embedder.embed_documents(chunks).   ,    ,        .

         ,   .

 

    sample_docs/   :

 test.txt   .

 test.pdf (   Word    PDF).

 test.docx   .

  :

python

# test_indexer.py

from embedder import Embedder

from vector_store import VectorStore

from indexer import index_folder

# 1.  

print(" embedding-...")

emb = Embedder()

store = VectorStore(collection_name="test_docs")

# 2.  

print(" ...")

count = index_folder(

folder_path="./sample_docs",

store=store,

embedder=emb,

chunk_size=500,

chunk_overlap=50

)

print(f"\n .  : {count}")

# 3.  -

query = "  "

query_vec = emb.embed_query(query)

results = store.search(query_vec, n_results=3)

print("\n :")

for doc, meta in zip(results["documents"][0], results["metadatas"][0]):

print(f"- [{meta['file_name']}] {doc[:100]}...")

:

bash

python test_indexer.py

:

text

 embedding- intfloat/multilingual-e5-base...

 ,  : 768

 ...

: ./sample_docs/test.txt

->  : 2

: ./sample_docs/test.docx

->  : 3

: ./sample_docs/test.pdf

->  : 5

 .  : 10

 :

- [test.pdf]  ... ,  ...

   

 ,   ?       ,         .    :    index_folder   .   ,      :

python

store.clear()

,  ,        / .            ,       .




    


PDF   .     ,   ,  PyPDF2             .        OCR, ,   Tesseract.      ,       ,  pytesseract    read_pdf.

  PDF   .           .    PDF    PyPDF2.PdfReader    ,    .     :      ,    ,   .

    .   TXT-   UTF-8.        Windows-1251,    .   read_txt   fallback:    UTF-8  ,  cp1251.        ,    chardet    .

     . RecursiveCharacterTextSplitter        , , .      ,        .    ,   chunk_overlap  ,  150  250 .         ,      .

 

   .          .   4.5   "  "      ask_documents(),   ,   ,           .      RAG.




 4.5.       


   ,     embedding-,   ,      .       ask_documents(),   ,        ,           .

      .      : ,  ,  .   ,      ,   .

:    ask_documents()

   :

1.     .          embedding-,     .

2.      ChromaDB.   , , 45     .

3.     .    ,          .

4.     .      ,        .

   ,    Python,    .

   RAG

        .    :

     .

  ,     .

  .

 ,    :

text

  -,    ,      .

:

1.      ,      .

2.          (: _).

3.        ,   .

4.       , : "     ."

5.   ,    .

:  ask_documents  LocalModel

      LocalModel.  local_model.py   :

python

# local_model.py (   )

from vector_store import VectorStore

from embedder import Embedder

class LocalModel:

# ... (    ) ...

def ask_documents(

self,

question: str,

store: VectorStore,

embedder: Embedder,

n_results: int = 4,

temperature: float = 0.3,

max_tokens: int = 512

) -> str:

"""

    .

:param question:  

:param store:  VectorStore   

:param embedder:  Embedder   

:param n_results:      

:param temperature:   (    )

:param max_tokens:   

:return:      

"""

#  1:  

query_embedding = embedder.embed_query(question)

#  2:   

results = store.search(query_embedding, n_results=n_results)

if not results["documents"] or not results["documents"][0]:

return "     ."

#  3:    

context_parts = []

for doc, meta in zip(results["documents"][0], results["metadatas"][0]):

file_name = meta.get("file_name", " ")

context_parts.append(f"---   : {file_name} ---\n{doc}\n")

context = "\n".join(context_parts)

system_prompt = (

"  -,    ,    "

"  . :\n"

"1.      ,      .\n"

"2.          (: _).\n"

"3.        ,   .\n"

"4.       , : \"     .\"\n"

"5.   ,    ."

)

user_message = (

f" : {question}\n\n"

f"   :\n\n{context}\n"

)

#  4:  

response = self.chat(

messages=[{"role": "user", "content": user_message}],

system_prompt=system_prompt,

temperature=temperature,

max_tokens=max_tokens

)

return response

 :

 temperature=0.3            .     .

    ,      ,         .

  

  test_rag.py,    :    .

python

# test_rag.py

from model_loader import download_model

from local_model import LocalModel

from embedder import Embedder

from vector_store import VectorStore

from indexer import index_folder

# 1. :  , embedding  

print("  ...")

model_path = download_model()

llm = LocalModel(model_path=model_path, n_threads=4, n_gpu_layers=0)

print(" embedding-...")

emb = Embedder()

print("  ...")

store = VectorStore(collection_name="my_docs")

# 2.  (      )

#         

store.clear()

count = index_folder(

folder_path="./sample_docs",

store=store,

embedder=emb,

chunk_size=500,

chunk_overlap=50

)

print(f" : {count}\n")

# 3.  

questions = [

"    ?",

"       ?",

"     ?",

"      ?"

]

for q in questions:

print(f"\n{'='*60}")

print(f": {q}")

print(f"{'='*60}")

answer = llm.ask_documents(

question=q,

store=store,

embedder=emb,

n_results=3,

temperature=0.2

)

print(f": {answer}")

:

bash

python test_rag.py

 :

text

  ...

 .

 embedding-...

 ,  : 768

  ...

: ./sample_docs/contract.txt

->  : 4

 : 4

============================================================

:     ?

============================================================

:   (: contract.txt),   

   30      .

      12  (: contract.txt).

============================================================

:        ?

============================================================

:         

 (: contract.txt).       

      30  (: contract.txt).

============================================================

:      ?

============================================================

:          0.1% 

  (: contract.txt).

============================================================

:       ?

============================================================

:      .

      ,   .    ,   .

  

  GUI     ask_documents(),      ,   .  :

python

def ask_documents_stream(

self,

question: str,

store: VectorStore,

embedder: Embedder,

n_results: int = 4,

temperature: float = 0.3,

max_tokens: int = 512

):

"""  ask_documents()."""

query_embedding = embedder.embed_query(question)

results = store.search(query_embedding, n_results=n_results)



if not results["documents"] or not results["documents"][0]:

yield "     ."

return

context_parts = []

for doc, meta in zip(results["documents"][0], results["metadatas"][0]):

file_name = meta.get("file_name", " ")

context_parts.append(f"---   : {file_name} ---\n{doc}\n")

context = "\n".join(context_parts)

system_prompt = (

"  -... : ..."  #   

)

user_message = f" : {question}\n\n :\n\n{context}"

yield from self.stream_chat(

messages=[{"role": "user", "content": user_message}],

system_prompt=system_prompt,

temperature=temperature,

max_tokens=max_tokens

)

   -

1.    ׸   .  Llama (  )   ,     .   ( )  .

2.     .  ,   , temperature=0.2..0.3   ,  0.7.   .

3.     .        .  3            .

4.     .        ,      n_ctx.  n_results=4  chunk_size=500   2000        2048     .     4096   n_results.




     RAG.


    .    ,   max_tokens   .    512  1024   2048   .          :  ,    ,    .  Llama  Qwen    .

    .           .  temperature  0.1  0.2             .    ,  :    ,    .       .       .

  ,    .                   .     ,     ,  Qwen 2.5 7B  Qwen 3 14B.     Llama,      :     ,   .     8B-.

    . ,      n_results     4  8,    .    ,     embedding-:      .           .         ,    .

 ,    .  ,     ,   file_name,        .        ,   indexer  relative_path      .         contract.pdf,  /2024/contract.pdf,   .

 

   4.  -  :

    .

       ChromaDB.

   ,      .

  ,   .

  .

  5    :     ,   .




 5.  





 5.1. - : Whisper tiny/base


          ,           .     :   .      ,          .      ,    Whisper,      .

  Whisper     

Whisper (https://github.com/openai/whisper)     ,  OpenAI     .    680 000             .         .

Whisper    :   tiny (39  ,  75   )   large-v3 (1.5  ,  3 ).       tiny  base. :

 tiny   ,         .          .

 base     ,   .    ,         .

    ,      .

 faster-whisper  whisper.cpp

 Whisper  OpenAI   Python,     .  ,    :

 faster-whisper    C++  (CTranslate2),    .    Python,        .

 whisper.cpp    C++   Python.  ,    ,    .     Python-  .

  faster-whisper.       :   pip,    Python-     .

 faster-whisper    

  :

bash

pip install faster-whisper sounddevice numpy

 faster-whisper    .

 sounddevice         (  Windows, macOS, Linux   ).

 numpy     .

 Linux      portaudio:

bash

sudo apt install portaudio19-dev # Debian/Ubuntu

   

  speech_recognition.py.  recognize_speech()          .

python

# speech_recognition.py

import queue

import numpy as np

import sounddevice as sd

from faster_whisper import WhisperModel

#  

SAMPLE_RATE = 16000          #   (  Whisper)

SILENCE_THRESHOLD = 0.01     #  ,    

SILENCE_DURATION = 1.5       #    ,   

MAX_RECORD_SECONDS = 30      #   

class SpeechRecognizer:

"""

     faster-whisper.

"""

def __init__(self, model_size: str = "tiny", device: str = "cpu"):

"""

:param model_size:  : "tiny", "base", "small", "medium", "large-v3"

:param device: "cpu"  "cuda" (  GPU NVIDIA)

"""

print(f"  Whisper ({model_size})...")

self.model = WhisperModel(model_size, device=device, compute_type="int8")

print(" .")

def record_until_silence(self) -> np.ndarray:

"""

      SILENCE_DURATION  

  MAX_RECORD_SECONDS.   float32   16 .

"""

q = queue.Queue()

recorded_chunks = []

def callback(indata, frames, time, status):

if status:

print(f": {status}")

q.put(indata.copy())

print("... ()")

with sd.InputStream(

samplerate=SAMPLE_RATE,

channels=1,

callback=callback,

dtype=np.float32

):

silence_start = None

while True:

chunk = q.get()

recorded_chunks.append(chunk)

#    

volume = np.abs(chunk).mean()

#   

if volume < SILENCE_THRESHOLD:

if silence_start is None:

silence_start = len(recorded_chunks) * chunk.shape[0] / SAMPLE_RATE

else:

silence_dur = (len(recorded_chunks) * chunk.shape[0] / SAMPLE_RATE) - silence_start

if silence_dur >= SILENCE_DURATION:

print(",  .")

break

else:

silence_start = None  # ,   

#   

total_sec = len(recorded_chunks) * chunk.shape[0] / SAMPLE_RATE

if total_sec >= MAX_RECORD_SECONDS:

print("  .")

break

#     

audio = np.concatenate(recorded_chunks, axis=0).flatten()

return audio

def transcribe(self, audio: np.ndarray) -> str:

"""

      Whisper.

"""

segments, info = self.model.transcribe(audio, language="ru", beam_size=5)

text = " ".join([seg.text for seg in segments])

return text.strip()

def recognize_speech(self) -> str:

"""

 :      .

      .

"""

audio = self.record_until_silence()

if audio is None or len(audio) == 0:

return ""

text = self.transcribe(audio)

return text

  :

  : Whisper     16   .   sd.InputStream .

  :      .     SILENCE_THRESHOLD,  .    SILENCE_DURATION  ,  .     ,  .

  : MAX_RECORD_SECONDS   ,      .

  language="ru"     .      ,   language=None     ,    .

 

  :

python

# test_speech.py

from speech_recognition import SpeechRecognizer

#     tiny

rec = SpeechRecognizer(model_size="tiny")

print(" Enter,   ...")

input()

text = rec.recognize_speech()

if text:

print(f": {text}")

else:

print("   .")

print("   .")

,  Enter    .      ,    .




   Whisper


       .            .

Tiny.         2 ,  3     .   75   .     :          .      ,   ,    .

Base.      3 ,      0.8 .   150 .    ,    :        .      .

Small.         5 ,   2 .   500 .    :           .           base     .

 :   tiny         .    ,   base.  small                    .               SpeechRecognizer.

     tiny  ,  base           .      : SpeechRecognizer(model_size="base").

   

 recognize_speech()     GUI  :

python

recognizer = SpeechRecognizer(model_size="tiny")

text = recognizer.recognize_speech()

if text:

#  text  LocalModel.chat()  ask_documents()

    recognize_speech   LocalModel,    . LocalModel   , SpeechRecognizer  .   5.3       .




    


 PortAudioError: ErroropeningInputStream.          ,    .     :      ,      .  Linux         pulseaudio --start   pipewire.    .

    . ,    transcribe   language="ru"    Whisper       ,    .  tiny  base  ,           .    ,   base  tiny   .

   .     ,    .  ,      .  SILENCE_DURATION  1.5  2.0          . ,  SILENCE_THRESHOLD  0.005          .

  .          .      tiny  base,    ,    .      NVIDIA,  device="cuda"  compute_type="float16"   SpeechRecognizer      .

  .          .         ,    dramatically.      ,        noisereduce,       .

 

    .            .  5.2  : piper-tts  silero ,      ,   ,    .




 5.2.  : piper-tts  silero.


   - .     .    ,     ,       ,        .      (Text-to-Speech, TTS),      ,     .  ,     .

  2026             ,    : piper-tts  silero.      ,      .

 piper-tts  silero

Piper-tts

Piper (https://github.com/rhasspy/piper)      ,   C++  Python-.        ,      Raspberry Pi. Piper  ,    ,  .

:

  ,       CPU.

   .

     (, ru_RU-ruslan-medium).

:

     (libportaudio, espeak-ng)    .

 Python- (piper-tts)     Windows.

    :    ,     .

Silero

Silero (https://github.com/snakers4/silero-models)         ,   .  TTS  Silero   -            .

:

    : , ,  .    (  ).

  :     PyTorch Hub  ,     .

    :    Python.  torch   .

  :  ,     .

:

  PyTorch,    200      (    LLM    ).

   piper-tts  CPU (     ).

  

  Silero.        piper-tts,    (    )      . PyTorch            ,          .

      (,      ),     Silero  Piper,    .   ,    .



  PyTorch     .  :

bash

pip install torch sounddevice

 Silero     ,      .

   

  speech_synthesis.py.     TextToSpeech,    Silero    .

python

# speech_synthesis.py

import os

import torch

import sounddevice as sd

class TextToSpeech:

"""

      Silero.

"""

def __init__(self, speaker: str = "xenia", device: str = "cpu"):

"""

:param speaker:  (      )

:param device: "cpu"  "cuda"

"""

self.device = torch.device(device)

print("  Silero TTS...")

#        PyTorch Hub

self.model, self.symbols, self.sample_rate, _, self.apply_tts = torch.hub.load(

repo_or_dir='snakers4/silero-models',

model='silero_tts',

language='ru',

speaker='ru_v3'

)

self.model = self.model.to(self.device)

self.speaker = speaker

print(f" .  : {self.speaker}")

def speak(self, text: str):

"""

       .

"""

if not text.strip():

return

# :    

audio = self.apply_tts(

texts=[text],

model=self.model,

symbol=self.symbols,

device=self.device,

speaker=self.speaker

)[0]  #   ( )  

#    numpy- float32

audio_np = audio.cpu().numpy()

# 

sd.play(audio_np, samplerate=self.sample_rate)

sd.wait()  #   

:

  : torch.hub.load   Silero (   )     :  ,  ,   ( 22050    ),      apply_tts.             .

  :  speaker   : 'xenia' (, ), 'aidar' (), 'baya' (,  ), 'kseniya' (), 'eugene' ().      'xenia'.     ,    Silero     .

 : sounddevice   float32      . sd.wait()     ,       .



python

# test_tts.py

from speech_synthesis import TextToSpeech

tts = TextToSpeech(speaker="xenia")

tts.speak("!    .   ?")

       .

  

  GUI    tts.speak(response)      .   (  ,      )     ,       .

   

 set_speaker     :

python

def set_speaker(self, speaker: str):

""" ."""

if speaker in self.model.speakers:  #   silero   

self.speaker = speaker

else:

print(f" {speaker}  ,  {self.speaker}")

          .




    


 ModuleNotFoundError: No module named 'torch'.  Silero   PyTorch,    .  pip install torch      CPU-,     .    GPU-,    CUDA- PyTorch .

  : HTTPError: Not Found.     TextToSpeech Silero     .         .   ,   ru_v3.pt    models.silero.ai (https://models.silero.ai/)       PyTorch Hub.      .

     .  ,    . Silero    22050   ,  sounddevice.play()     .     ,    samplerate=22050   sd.play().         .

      . Silero         . ,      ,      privet.  ,      , xenia  aidar.        .

 PortAudioError  .       . ,           .  Linux       pulseaudio --start.   ,            .

 : piper-tts (   )

   -     PyTorch,     piper-tts:

bash

#  piper-tts

pip install piper-tts

#  ru_RU-ruslan-medium  GitHub  models/

python

from piper import PiperVoice

import sounddevice as sd

import wave

voice = PiperVoice.load("models/ru_RU-ruslan-medium.onnx")

audio = voice.synthesize(", !")

# audio   bytes  WAV,    

     Silero:        200  ,       .

 

      ( ),   ().    5.3            LocalModel     ,    ,           .




 5.3.    .


   ,   -     .   5.1           Whisper.   5.2     Silero TTS.          LocalModel    ,   ,    .      ,     (   ).

  

     Python- VoiceAssistant,    :

1.    SpeechRecognizer        .

2.    LocalModel   ,   (   RAG,   VectorStore).

3.    TextToSpeech       .

        .     ,     .

 VoiceAssistant

  voice_assistant.py:

python

# voice_assistant.py

import sys

from speech_recognition import SpeechRecognizer

from speech_synthesis import TextToSpeech

from local_model import LocalModel

from vector_store import VectorStore

from embedder import Embedder

class VoiceAssistant:

"""

 ,   ,

    .

"""

def __init__(

self,

model: LocalModel,

whisper_size: str = "tiny",

tts_speaker: str = "xenia",

store: VectorStore = None,

embedder: Embedder = None,

use_documents: bool = False

):

"""

:param model:  LocalModel

:param whisper_size:   Whisper ("tiny" / "base")

:param tts_speaker:  Silero TTS

:param store: VectorStore (  RAG)

:param embedder: Embedder (  RAG)

:param use_documents:  " "  

"""

self.model = model

self.store = store

self.embedder = embedder

self.use_documents = use_documents

print("  ...")

self.recognizer = SpeechRecognizer(model_size=whisper_size)

print("  ...")

self.tts = TextToSpeech(speaker=tts_speaker)

print("  .\n")

def listen(self) -> str:

"""     ."""

text = self.recognizer.recognize_speech()

if text:

print(f": {text}")

return text

def respond(self, text: str):

"""    ."""

print("...", end="", flush=True)

# ,    

if self.use_documents and self.store and self.embedder:

response = self.model.ask_documents(

question=text,

store=self.store,

embedder=self.embedder

)

else:

response = self.model.chat(

messages=[{"role": "user", "content": text}],

temperature=0.7

)

print(f"\r: {response}")

#  

self.tts.speak(response)

def run(self, wake_word: str = None):

"""

 .  wake_word ,    .

      Enter.

"""

print("=" * 50)

print("  .")

print(" Ctrl+C  .\n")

if wake_word:

print(f"   : \"{wake_word}\"...")

else:

print(" Enter,   ...")

try:

while True:

if wake_word:

#     

text = self.listen()

if text and wake_word.lower() in text.lower():

print(f"  !")

#     

query = text.lower().replace(wake_word.lower(), "").strip()

if query:

self.respond(query)

else:

#      ,  

self.tts.speak(" .")

else:

#  :  Enter

input()

text = self.listen()

if text:

#    

if text.lower() in ["", "", ""]:

self.tts.speak(" !")

break

elif text.lower() in [" "]:

self.use_documents = not self.use_documents

status = "" if self.use_documents else ""

self.tts.speak(f"  {status}")

continue

else:

self.respond(text)

except KeyboardInterrupt:

print("\n ...")

self.tts.speak(" .")

  :

     .  store  embedder :      RAG,    .  use_documents     .

  listen()    SpeechRecognizer.recognize_speech(),    .

  respond()         ,  ask_documents().        .      \r  end="":   ...,    ,           :.

  run()   .   :

o  ( Enter):      ,       ,  .

o   :            (, ).       .       ,     .

     :

 , ,    .

          RAG.

  

  test_voice_assistant.py:

python

# test_voice_assistant.py

from model_loader import download_model

from local_model import LocalModel

from voice_assistant import VoiceAssistant

from vector_store import VectorStore

from embedder import Embedder

#   

print("  ...")

model_path = download_model()

llm = LocalModel(model_path=model_path, n_threads=4)

#     ,  

store = VectorStore(collection_name="my_docs")

emb = Embedder()

use_docs = True  #  False,   

#  

assistant = VoiceAssistant(

model=llm,

whisper_size="tiny",

tts_speaker="xenia",

store=store,

embedder=emb,

use_documents=use_docs

)

# 

assistant.run()  #   ( Enter)

:

bash

python test_voice_assistant.py

  ( ):

text

  ...

 .

  ...

 .

  .

==================================================

  .

 Ctrl+C  .

 Enter,   ...

[Enter]

... ()

,  .

:    

...

:  ,       .

( )

[Enter]

... ()

:    

:      ?   31 OCT = 25 DEC!

()

   

  ,         ,  :

python

assistant.run(wake_word="")

       ,    ,   ?. ,       .         ,        .

  ()

     ,    .     ,      .       ,       .   :

     ( ,    ).

   ,    .

   ,       .     respond()   .




    


      .  ,   SILENCE_THRESHOLD,     .            0.005   0.003. , ,     -       0.02.     :    ,      .

Whisper    .         ,    .    (Mic Boost),    .      ,   tiny  base           ,    .




  .


   .

   ,     (https://www.litres.ru/book/valeriy-antonov-3298/chatgpt-na-vashem-noutbuke-besplatno-anonimno-navsegd-74131963/)  .

      Visa, MasterCard, Maestro,    ,   ,     ,  PayPal, WebMoney, ., QIWI ,       .


