一键重装系统工具 | U盘启动盘制作工具 | 误删文件恢复软件 | 硬盘数据抢救专家 | 电脑蓝屏修复助手 | C盘空间清理神器 | 电脑驱动离线安装工具 | 微信聊天记录恢复工具 | 照片误格式化恢复 | 电脑密码破解清除工具 | 系统崩溃紧急救援盘 | 电脑加速优化大师 | 电脑开不了机怎么重装系统 | 回收站清空了怎么恢复 | 硬盘分区丢失数据恢复 | 电脑卡顿重装系统有用吗 | U盘插入提示格式化数据恢复 | 电脑中毒文件被隐藏恢复 | 忘记电脑开机密码怎么办 | 新硬盘分区对齐工具 | 旧电脑装Win10流畅工具 | SD卡照片删除恢复免费版 | 移动硬盘打不开提示损坏修复 | 电脑无故重启系统修复工具 | 电脑小白一键重装神器 | 程序员电脑环境配置助手 | 设计师电脑字体/素材恢复工具 | 网吧网管系统维护工具箱 | 财务人员电脑发票备份恢复 | 学生党免费电脑系统安装包 | 电脑维修师傅必备工具盘 | 游戏玩家电脑性能优化助手 | 办公白领误删文档恢复软件 | 自媒体视频素材恢复工具 | 网课录制视频损坏修复工具 | 最好的U盘PE系统排名 | 数据恢复软件哪个最强 | 免费电脑助手与收费版区别 | 国产装机工具哪款无广告 | 离线版驱动助手推荐 | 轻量级电脑优化工具对比 | 支持NVMe驱动的PE工具 | 带网络功能的应急启动盘 | 2026最新版万能装机工具 | 支持Win11 24H2的PE工具 | 最新免激活系统重装工具 | 2026数据恢复软件破解版合集 | 纯净无捆绑装机助手V3.0 | 支持苹果M芯片的电脑助手 | 秋季更新版系统维护工具箱 | 电脑系统崩了怎么用U盘把重要资料拷贝出来 | 重装系统前哪些文件夹必须备份 | 固态硬盘误格式化还能恢复数据吗 | 如何制作一个既带PE又能存数据的双分区U盘 | 电脑总是弹窗广告用什么助手彻底拦截 后台管理
📢 欢迎访问系统之家!所有资源均经过安全检测。

Free Browser Based AI LLM вҖ” Run an LLM locally in your browser

发布时间:2026-09-21 | 浏览:1
📥 下载地址(文章开头)
装机神器,在线重装利器,在线安装一切系统。
How this works е·ҘдҪңеҺҹзҗҶ Model comparison жЁЎеһӢеҜ№жҜ” Why run an LLM in your browser? дёәд»Җд№ҲиҰҒеңЁжөҸи§ҲеҷЁдёӯиҝҗиЎҢ LLM? Most free AI chatbots send every prompt to a cloud server, require an account, and log your conversations. This tool works differently: it downloads an open-weights language model вҖ” Llama, Qwen, Phi, Gemma or Mistral вҖ” straight into your browser tab and runs it on your own hardware via WebGPU, or on any CPU through WebAssembly. No signup, no API key, no subscription, and no data collection of any kind. еӨ§еӨҡж•°е…Қиҙ№ AI иҒҠеӨ©жңәеҷЁдәәдјҡжҠҠжҜҸдёҖжқЎжҸҗзӨәеҸ‘йҖҒеҲ°дә‘з«ҜжңҚеҠЎеҷЁ,йңҖиҰҒжіЁеҶҢиҙҰжҲ·,е№¶и®°еҪ•дҪ зҡ„еҜ№иҜқгҖӮиҝҷдёӘе·Ҙе…·зҡ„еҒҡжі•дёҚеҗҢ:е®ғдјҡжҠҠдёҖдёӘејҖж”ҫжқғйҮҚзҡ„иҜӯиЁҖжЁЎеһӢвҖ”вҖ”LlamaгҖҒQwenгҖҒPhiгҖҒGemma жҲ– MistralвҖ”вҖ”зӣҙжҺҘдёӢиҪҪеҲ°дҪ зҡ„жөҸи§ҲеҷЁж ҮзӯҫйЎөдёӯ,е№¶йҖҡиҝҮ WebGPU еңЁдҪ иҮӘе·ұзҡ„зЎ¬д»¶дёҠиҝҗиЎҢ,жҲ–иҖ…йҖҡиҝҮ WebAssembly еңЁд»»ж„Ҹ CPU дёҠиҝҗиЎҢгҖӮж— йңҖжіЁеҶҢгҖҒж— йңҖ API keyгҖҒж— йңҖи®ўйҳ…,д№ҹдёҚж”¶йӣҶд»»дҪ•ж•°жҚ®гҖӮ Private by design вҖ” your prompts never leave your device, which makes this safe for drafts, client notes, or anything you wouldn't paste into a cloud chatbot. Works offline вҖ” after the one-time model download you can keep chatting on a plane or behind a strict firewall. Free without limits вҖ” no trial, no message caps, no upsell; the models are open-source and the compute is yours. и®ҫи®ЎдёҠе°ұжҳҜз§ҒеҜҶзҡ„ вҖ”вҖ” дҪ зҡ„жҸҗзӨәиҜҚж°ёиҝңдёҚдјҡзҰ»ејҖдҪ зҡ„и®ҫеӨҮ,иҝҷдҪҝеҫ—е®ғеҸҜд»Ҙе®үе…Ёең°з”ЁдәҺиҚүзЁҝгҖҒе®ўжҲ·з¬”и®°,жҲ–иҖ…д»»дҪ•дҪ дёҚж„ҝзІҳиҙҙиҝӣдә‘з«ҜиҒҠеӨ©жңәеҷЁдәәзҡ„еҶ…е®№гҖӮ зҰ»зәҝеҸҜз”Ё вҖ”вҖ” дёҖж¬ЎжҖ§дёӢиҪҪжЁЎеһӢд№ӢеҗҺ,дҪ еҸҜд»ҘеңЁйЈһжңәдёҠжҲ–дёҘж јзҡ„йҳІзҒ«еўҷеҗҺйқўз»§з»ӯиҒҠеӨ©гҖӮ е…Қиҙ№дё”ж— йҷҗеҲ¶ вҖ”вҖ” жІЎжңүиҜ•з”ЁжңҹгҖҒжІЎжңүж¶ҲжҒҜдёҠйҷҗгҖҒжІЎжңүд»ҳиҙ№еҚҮзә§;жЁЎеһӢжҳҜејҖжәҗзҡ„,з®—еҠӣжҳҜдҪ иҮӘе·ұзҡ„гҖӮ Typical uses: rewriting and summarizing text, drafting emails, explaining code, generating KQL or SQL queries, extracting JSON from messy text, brainstorming, and language practice вҖ” a lightweight, private alternative to cloud AI assistants, running as a local AI chatbot in your browser. е…ёеһӢз”ЁйҖ”еҢ…жӢ¬:ж”№еҶҷе’ҢжҖ»з»“ж–Үжң¬гҖҒиө·иҚүйӮ®д»¶гҖҒи§ЈйҮҠд»Јз ҒгҖҒз”ҹжҲҗ KQL жҲ– SQL жҹҘиҜўгҖҒд»ҺжқӮд№ұж–Үжң¬дёӯжҸҗеҸ– JSONгҖҒеӨҙи„‘йЈҺжҡҙд»ҘеҸҠиҜӯиЁҖз»ғд№ вҖ”вҖ” дҪңдёәдә‘з«Ҝ AI еҠ©жүӢзҡ„иҪ»йҮҸзә§гҖҒз§ҒеҜҶжӣҝд»Јж–№жЎҲ,д»Ҙжң¬ең° AI иҒҠеӨ©жңәеҷЁдәәзҡ„еҪўејҸиҝҗиЎҢеңЁдҪ зҡ„жөҸи§ҲеҷЁдёӯгҖӮ More local-AI tools on this site: IBM Granite AI in your browser В· Whisper speech-to-text (local) В· LLM token counter В· AI-generated text detector жң¬з«ҷжӣҙеӨҡжң¬ең° AI е·Ҙе…·: IBM Granite AI жөҸи§ҲеҷЁзүҲ В· Whisper жң¬ең°иҜӯйҹіиҪ¬еҶҷ В· LLM token и®Ўж•°еҷЁ В· AI з”ҹжҲҗж–Үжң¬жЈҖжөӢеҷЁ 01. Getting started A real language model running inside your browser tab — not a front-end for someone else's API. You pick an open-weights model, it downloads once, and from then on every prompt is processed on your own GPU or CPU. There is no account, no server doing the thinking, and no per-message cost. Open the page, pick a model and press Load model . The weights download once, are cached in your browser, and then run locally — on your GPU through WebGPU, or on your CPU in the Mobile / WASM tab. There is no install, no command line and no server involved. No app, no runtime, no admin rights. The only thing that downloads is the model file itself, straight into the browser cache. Nothing is written to your Downloads folder and nothing is registered with the operating system — clearing site data removes every trace. None of the three. Because the model runs on your own hardware there is nothing to authenticate against and no per-token bill. You never enter an email address, a key or a card number. You are downloading the model itself — anything from ~70 MB for the smallest mobile model to several gigabytes for a 7B one. It is stored with the browser's Cache API, so the second visit starts in seconds and needs no network at all. A cached model is marked with a badge in the picker. Yes, once the weights are cached. You can switch off Wi-Fi entirely and keep chatting — a useful test in itself, because a tool that still answers with the network disconnected clearly is not calling a server. Only the first download of each model needs a connection. It asks your GPU what the largest buffer it can allocate is, and picks a model that comfortably fits: a 1B model on modest hardware, Phi 3.5 Mini on a mid-range GPU, a 7B model on a card with several gigabytes to spare. It is a starting point rather than a verdict — you can always pick something bigger or smaller by hand. Completely free, with no account, no subscription, no message cap and no watermark. There is no rate limit either, because there is no server to rate-limit — the only ceiling is how fast your own hardware can generate tokens. 02. Choosing a model Llama 3.2 1B is the sensible default: around 880 MB, quick to load, and reliable at rewriting, summarising and everyday instructions. Qwen 2.5 0.5B (~350 MB) is the one to choose when speed matters more than depth. Phi 3.5 Mini (~2.2 GB) gives the best reasoning, coding and structured output if your GPU can hold it, and Gemma 2 2B writes the most natural prose. Start at 1B and move up only if the answers disappoint you. The number of parameters, in billions — roughly, how much the model knows and how well it can reason. Each step up costs download size, memory and speed: a 1B model is around 880 MB and feels instant, a 3B is about 2 GB and noticeably more coherent, a 7B is over 4 GB and needs a proper graphics card. Quality rises with size, but so does the wait. Qwen 2.5 Coder 1.5B is trained specifically on code and is remarkably good for its size — the best value if you want completions, small functions and shell one-liners. Phi 3.5 Mini is the stronger all-rounder when you need the model to reason about the code rather than just produce it. Neither will replace a frontier model on a large codebase. Small models trained on the output of a much larger reasoning model, so they work through a problem step by step before answering. That makes them better at puzzles and multi-step logic than their size suggests — and slower, because they spend tokens thinking out loud. Raise max tokens when you use them or they will be cut off mid-thought. The Qwen 2.5 family is the strongest multilingual option here, with 1.5B and 3B versions covering a wide range of languages. Phi 3.5 Mini also handles 20+ languages well. Small models drift more in languages other than English, so keep prompts explicit — setting the system prompt to “always answer in Dutch” works better than asking mid-conversation. Only with a discrete GPU that has roughly 5 GB or more of free video memory — the weights alone are around 4.3 GB before the conversation cache. Mistral 7B and the DeepSeek R1 7B distill are in the picker for machines that can take them. On integrated graphics they will either fail to allocate or crawl, and a 3B model is the better trade. Much smaller ones, because it runs on the CPU: SmolLM2 135M (~70 MB), SmolLM2 360M (~200 MB, the recommended default), Qwen 1.5 0.5B and SmolLM2 1.7B for devices with memory to spare. They are noticeably simpler than the desktop models, but they run anywhere — no WebGPU required. Each model is cached separately, so the first switch to a new model downloads it and every later switch is instant. Models you have already cached are flagged in the picker, which makes it easy to hop between a fast one and a smarter one without paying the download twice. It is how much text the model can hold in mind at once — your system prompt, the conversation so far and its own reply all count against it. Phi 3.5 Mini is the roomiest here at 128K tokens; older models such as Phi 3 Mini 4K are limited to a few thousand. When a long chat starts to lose the plot, reset it and paste in only what matters. It is the quantisation: the weights are compressed to 4 bits with 16-bit floating point used during computation. That is what makes a multi-billion-parameter model small enough to download and fit in browser memory. The cost is a small amount of accuracy, which is a bargain compared with not being able to run the model at all. 03. Privacy and data No. Prompts and replies never leave the tab. The single network request the tool makes is the one-time download of the model weights from the Hugging Face CDN; after that you can disconnect entirely and it keeps working. There is no inference server, because your device is the inference server. The conversation lives in the page and disappears when you close or reset the tab. Nothing is uploaded, nothing is written to a database, and the only things kept in browser storage are your theme choice and a note of which models you have already cached — never the messages. Use copy log or ↓ .md if you want to keep a conversation. Genuinely, and you can check it yourself rather than take anyone's word for it. There is no API key in the page to send anything with, no endpoint to send it to, and the model answers with the network switched off — which is impossible for a cloud chatbot. Two ways. Open your browser's developer tools, go to the Network tab, and watch it while you send a message: after the model has loaded there are no requests at all. Or simply disable Wi-Fi and keep chatting — if the answers still come, they are being generated locally. Technically nothing leaves your device, so there is no transfer to a third party and no processor to sign an agreement with — which is exactly why local models are attractive for sensitive text. That said, your own organisation's policy still applies, so check it before pasting client data anywhere, including here. Downloading the weights is an ordinary file request to a CDN, so that CDN sees a file being fetched, as it would for any image or script. It carries no prompt, no conversation and no identity beyond a normal web request — and it happens once per model, not per message. It works, but private windows discard storage when you close them, so the model downloads again the next time. If you plan to use it regularly, a normal window keeps the cache and saves you the wait. 04. Speed and hardware For the desktop tab, a GPU that supports WebGPU and has enough free memory for the model: roughly 1 GB for a 1B model, 2–3 GB for a 3B, and 5 GB or more for a 7B. Integrated graphics handle the 0.5B–1B models comfortably. If none of that applies to your machine, the Mobile / WASM tab runs on the CPU on practically anything. On a modern discrete GPU a 1B model typically streams faster than you can read, and a 3B model still feels conversational. On integrated graphics expect it to be readable but unhurried, and in CPU / WASM mode expect a few words per second. The live tok/s counter under the chat box shows exactly what your device is doing. Usually the model is too large for the hardware, so memory is being shuffled instead of used. Drop a size class, close other GPU-heavy tabs — video calls and 3D pages compete for the same memory — and check that your laptop is not in a battery-saver profile that caps the GPU. A long conversation also slows things down, because the whole history is reprocessed; reset the chat to get the speed back. It hit the max tokens limit, which defaults to 1024. For scripts, long code or detailed documents raise it to 2048, 4096 or 8192 next to the Send button. It is a deliberate cap rather than a failure — without it a small model can ramble for a very long time. It controls how adventurous the sampling is. Low values (around 0.2) make the model repetitive but predictable, which is what you want for extraction, JSON and code. Higher values (0.8 and above) produce more varied writing at the cost of accuracy. The default of 0.7 is a reasonable middle for chat. Typically several times faster, sometimes an order of magnitude, which is why the WebGPU tab is the default on desktop. WASM mode exists so that the tool still works on hardware and browsers without WebGPU — it trades speed for running absolutely anywhere. The model did not fit. Choose a smaller one, close other tabs using the GPU, and try again — the error message names the limit it hit. On laptops that switch between integrated and discrete graphics, forcing the browser onto the discrete GPU in the system settings often solves it outright. While it is generating, yes — you are running a neural network on your own silicon, so the fans may spin up much as they would during a game. It only draws power while a reply is being produced; the model sitting in memory idle costs nothing. 05. Browsers, storage and troubleshooting Chrome and Edge have shipped it by default since version 113 on Windows, macOS and ChromeOS, and on Android 12+ since Chrome 121. Safari enables it by default in macOS Tahoe 26, iOS 26 and iPadOS 26. Firefox ships it on Windows from version 141 and on Apple Silicon Macs from 147, with Linux and Android still in Nightly. Anything not on that list can still use the Mobile / WASM tab. Your browser, or that machine's graphics driver, is not exposing WebGPU — common on older browser versions, on Linux with certain drivers, in some virtual machines, and where enterprise policy disables it. The page detects this and moves you to the Mobile / WASM tab automatically, so you can still use the tool; updating the browser is the fix if you want GPU speed. Yes, in the Mobile / WASM tab, with one caveat: iOS Safari enforces a hard memory ceiling of roughly 1–1.5 GB per tab, so only the smallest model (SmolLM2 135M) loads reliably. Larger ones crash the tab with “a problem repeatedly occurred”. For anything more capable, open the page on a desktop. Yes. Chrome on Android 12 and later supports WebGPU on most recent chipsets, and where it does not, the WASM tab runs on the CPU. Stick to the small models — phone memory limits bite long before the model quality does. Corporate proxies and content filters frequently block the Hugging Face CDN, which is where the weights come from, and some inspect and break large downloads. That is a network policy question rather than a browser one — the same page usually works immediately on a home connection or a phone hotspot. In the browser's Cache Storage for this site, not in your file system. Clearing site data (or “Cookies and other site data” for this domain) removes every cached model and frees the space; the tool then behaves like a first visit. Nothing outside the browser profile is touched. Yes. Browsers evict cached data when disk space runs low or a profile is cleaned, and private windows discard it on close. If a model that loaded instantly yesterday starts downloading again, that is what happened — nothing is broken. Only what the models you actually load take up: from about 70 MB for the smallest to roughly 4.5 GB for a 7B. Every model you try is kept until you clear site data, so trying several does add up — worth remembering on a machine that is short on space. 06. What it is good at — and what it is not Because it is smaller by three orders of magnitude. Frontier models run on server racks with hundreds of billions of parameters; a 1B model in a browser tab is a fraction of a percent of that, and it is remarkable that it works at all. Judge it as a fast local assistant for small jobs, not as a replacement for a frontier model. Short, well-defined text work: rewriting a paragraph formally or casually, summarising something you paste in, drafting a reply, extracting fields into JSON, explaining a concept simply, and small code and query snippets. The preset buttons above the chat box are shortcuts to exactly those tasks. Anything that needs current facts or long reasoning chains is the wrong job for it. No. It has no internet access and no access to your files — it only sees what you type into the box. That is the same property that makes it private: there is no channel in either direction. If you want it to work on a document, paste the relevant part in. More often than a large model, yes. Small models are confident and frequently wrong about names, dates, numbers and anything specialised, and they have no way to look anything up. Treat factual claims as drafts to verify, and lean on the tasks where the source material is in your prompt — rewriting, summarising, extracting. Yes. Open System prompt (persona) under the model picker and describe how it should respond — “you are a terse senior engineer”, “always answer in Dutch”, “reply only with JSON”. It applies to every message in the conversation, and it is by far the most effective way to improve the output of a small model. For small, self-contained things, yes — a regex, a shell one-liner, a function, a KQL query from the preset. Use Qwen 2.5 Coder or Phi 3.5 Mini, keep the request narrow, raise max tokens so it is not truncated, and read what it produces before running it. It has no idea what is in your codebase. Generally yes, but the licence belongs to the model rather than to this page, and they differ: Llama has Meta's community licence, Gemma has Google's terms of use, most Qwen and SmolLM2 releases are Apache 2.0 and Phi is MIT. If the output is going into a product, read the licence for the specific model you used. 07. Under the hood
📥 下载地址(文章中间)
装机神器,在线重装利器,在线安装一切系统。
The desktop tab uses WebLLM from MLC AI, which compiles open-weights models to run on WebGPU . The mobile tab uses Hugging Face's transformers.js on top of ONNX Runtime Web, executing on the CPU through WebAssembly. Everything around them is plain JavaScript and CSS in a single HTML file. From the Hugging Face CDN, in the pre-compiled MLC format for the WebGPU tab and ONNX for the WASM tab. They are open-weights models published by Meta, Alibaba, Microsoft, Google DeepMind, Hugging Face and DeepSeek — the same files anyone can download and run locally. Same idea, different delivery. Those are desktop applications you install, and they will always be faster and support far larger models because they use your hardware directly. This runs the model in a browser tab with nothing installed at all, which makes it the easier way to try a local model — on a locked-down machine, on someone else's computer, or simply to see what the fuss is about. When the task needs breadth, accuracy on facts, long documents, up-to-date information or serious reasoning — a frontier model will do in one attempt what a 1B model cannot do at all. Use this when the text is sensitive, when you are offline, when you want an instant answer without opening an account, or when the job is small enough that a small model is simply the quicker tool. 01. еҝ«йҖҹдёҠжүӢ дёҖдёӘзңҹжӯЈзҡ„иҜӯиЁҖжЁЎеһӢиҝҗиЎҢеңЁ дҪ зҡ„жөҸи§ҲеҷЁж ҮзӯҫйЎөйҮҢ ——иҖҢдёҚжҳҜеҲ«дәә API зҡ„еүҚз«Ҝз•ҢйқўгҖӮдҪ йҖүжӢ©дёҖдёӘејҖж”ҫжқғйҮҚжЁЎеһӢ,дёӢиҪҪдёҖж¬Ў,жӯӨеҗҺжҜҸдёҖжқЎжҸҗзӨәйғҪеңЁдҪ иҮӘе·ұзҡ„ GPU жҲ– CPU дёҠеӨ„зҗҶгҖӮжІЎжңүиҙҰжҲ·гҖҒжІЎжңүжңҚеҠЎеҷЁжӣҝдҪ "жҖқиҖғ",д№ҹжІЎжңүжҢүж¶ҲжҒҜи®Ўиҙ№гҖӮ жү“ејҖйЎөйқў,йҖүжӢ©дёҖдёӘжЁЎеһӢ,зӮ№еҮ» "еҠ иҪҪжЁЎеһӢ" гҖӮжқғйҮҚеҸӘйңҖдёӢиҪҪдёҖж¬Ў,дјҡзј“еӯҳеңЁдҪ зҡ„жөҸи§ҲеҷЁдёӯ,д№ӢеҗҺе°ұеңЁжң¬ең°иҝҗиЎҢ——еңЁ Desktop ж ҮзӯҫйЎөйҖҡиҝҮ WebGPU дҪҝз”ЁдҪ зҡ„ GPU,жҲ–еңЁ Mobile / WASM ж ҮзӯҫйЎөдҪҝз”ЁдҪ зҡ„ CPUгҖӮе…ЁзЁӢ ж— йңҖе®үиЈ…гҖҒж— йңҖе‘Ҫд»ӨиЎҢгҖҒд№ҹдёҚж¶үеҸҠд»»дҪ•жңҚеҠЎеҷЁ гҖӮ дёҚйңҖиҰҒ AppгҖҒдёҚйңҖиҰҒиҝҗиЎҢж—¶гҖҒд№ҹдёҚйңҖиҰҒз®ЎзҗҶе‘ҳжқғйҷҗгҖӮе”ҜдёҖдјҡдёӢиҪҪзҡ„жҳҜжЁЎеһӢж–Үд»¶жң¬иә«,зӣҙжҺҘеӯҳе…ҘжөҸи§ҲеҷЁзј“еӯҳгҖӮдёҚдјҡеҶҷе…ҘдҪ зҡ„"дёӢиҪҪ"ж–Үд»¶еӨ№,д№ҹдёҚдјҡеңЁж“ҚдҪңзі»з»ҹдёӯжіЁеҶҢд»»дҪ•еҶ…е®№——жё…йҷӨзҪ‘з«ҷж•°жҚ®еҚіеҸҜжҠ№еҺ»жүҖжңүз—•иҝ№гҖӮ дёүиҖ…йғҪдёҚйңҖиҰҒгҖӮеӣ дёәжЁЎеһӢиҝҗиЎҢеңЁдҪ иҮӘе·ұзҡ„зЎ¬д»¶дёҠ,жІЎжңүд»»дҪ•йңҖиҰҒйӘҢиҜҒиә«д»Ҫзҡ„еҜ№иұЎ,д№ҹжІЎжңүжҢү token и®Ўиҙ№зҡ„иҙҰеҚ•гҖӮдҪ ж— йңҖиҫ“е…ҘйӮ®з®ұгҖҒеҜҶй’ҘжҲ–й“¶иЎҢеҚЎеҸ·гҖӮ дҪ жӯЈеңЁдёӢиҪҪжЁЎеһӢжң¬иә«——жңҖе°Ҹзҡ„з§»еҠЁз«ҜжЁЎеһӢзәҰ 70 MB,иҖҢ 7B жЁЎеһӢеҲҷжңүеҘҪеҮ GBгҖӮе®ғйҖҡиҝҮжөҸи§ҲеҷЁзҡ„ Cache API еӯҳеӮЁ,жүҖд»Ҙ第дәҢж¬Ўи®ҝй—®еҮ з§’й’ҹе°ұиғҪеҗҜеҠЁ,е®Ңе…ЁдёҚйңҖиҰҒиҒ”зҪ‘гҖӮе·Ізј“еӯҳзҡ„жЁЎеһӢдјҡеңЁйҖүжӢ©еҷЁдёӯж ҮеҮәеҫҪз« гҖӮ еҸҜд»Ҙ,еҸӘиҰҒжқғйҮҚе·Ізј“еӯҳгҖӮдҪ еҸҜд»Ҙе®Ңе…Ёе…ій—ӯ Wi-Fi з»§з»ӯиҒҠеӨ©——иҝҷжң¬иә«е°ұжҳҜдёҖдёӘеҫҲеҘҪзҡ„жөӢиҜ•,еӣ дёәж–ӯзҪ‘еҗҺд»ҚиғҪеӣһзӯ”й—®йўҳ,жҳҫз„¶иҜҙжҳҺе®ғжІЎжңүеңЁи°ғз”ЁжңҚеҠЎеҷЁгҖӮеҸӘжңүжҜҸдёӘжЁЎеһӢ第дёҖж¬ЎдёӢиҪҪж—¶йңҖиҰҒиҒ”зҪ‘гҖӮ е®ғдјҡиҜўй—®дҪ зҡ„ GPU иғҪеҲҶй…Қзҡ„жңҖеӨ§зј“еҶІеҢәжҳҜеӨҡе°‘,з„¶еҗҺжҢ‘йҖүдёҖдёӘиғҪиҲ’йҖӮиҝҗиЎҢзҡ„жЁЎеһӢ:жҷ®йҖҡзЎ¬д»¶й…Қ 1B жЁЎеһӢ,дёӯз«Ҝ GPU й…Қ Phi 3.5 Mini,жӢҘжңүе……иЈ•жҳҫеӯҳзҡ„жҳҫеҚЎй…Қ 7B жЁЎеһӢгҖӮиҝҷеҸӘжҳҜдёҖдёӘиө·зӮ№,иҖҢйқһе®ҡи®ә——дҪ йҡҸж—¶еҸҜд»ҘжүӢеҠЁйҖүжӢ©жӣҙеӨ§жҲ–жӣҙе°Ҹзҡ„жЁЎеһӢгҖӮ е®Ңе…Ёе…Қиҙ№,ж— йңҖиҙҰжҲ·гҖҒж— йңҖи®ўйҳ…гҖҒжІЎжңүж¶ҲжҒҜдёҠйҷҗ,д№ҹжІЎжңүж°ҙеҚ°гҖӮд№ҹжІЎжңүйҖҹзҺҮйҷҗеҲ¶,еӣ дёәж №жң¬жІЎжңүжңҚеҠЎеҷЁеҸҜдҫӣйҷҗйҖҹ——е”ҜдёҖзҡ„дёҠйҷҗжҳҜдҪ иҮӘе·ұзЎ¬д»¶з”ҹжҲҗ token зҡ„йҖҹеәҰгҖӮ 02. йҖүжӢ©жЁЎеһӢ Llama 3.2 1B жҳҜжҳҺжҷәзҡ„й»ҳи®ӨйҖүжӢ©:зәҰ 880 MB,еҠ иҪҪеҝ«,еңЁж”№еҶҷгҖҒжҖ»з»“е’Ңж—ҘеёёжҢҮд»ӨдёҠйғҪеҫҲеҸҜйқ гҖӮ Qwen 2.5 0.5B (зәҰ 350 MB)йҖӮеҗҲеңЁйҖҹеәҰжҜ”ж·ұеәҰжӣҙйҮҚиҰҒж—¶дҪҝз”ЁгҖӮ Phi 3.5 Mini (зәҰ 2.2 GB)еҰӮжһңдҪ зҡ„ GPU иғҪжүҝеҸ—,дјҡз»ҷеҮәжңҖдҪізҡ„жҺЁзҗҶгҖҒзј–зЁӢе’Ңз»“жһ„еҢ–иҫ“еҮәгҖӮиҖҢ Gemma 2 2B еҶҷеҮәзҡ„ж–Үеӯ—жңҖиҮӘз„¶гҖӮе…Ҳд»Һ 1B ејҖе§Ӣ,еҸӘжңүеңЁеӣһзӯ”дёҚиғҪи®©дҪ ж»Ўж„Ҹж—¶еҶҚеҫҖдёҠеҚҮзә§гҖӮ иҝҷжҳҜеҸӮж•°йҮҸ,д»ҘеҚҒдәҝ(B)дёәеҚ•дҪҚ——еӨ§иҮҙд»ЈиЎЁжЁЎеһӢзҹҘйҒ“еӨҡе°‘гҖҒжҺЁзҗҶиғҪеҠӣжңүеӨҡејәгҖӮжҜҸдёҠдёҖдёӘжЎЈж¬ЎйғҪдјҡеўһеҠ дёӢиҪҪдҪ“з§ҜгҖҒеҶ…еӯҳеҚ з”Ёе’ҢиҝҗиЎҢйҖҹеәҰзҡ„д»Јд»·:1B жЁЎеһӢзәҰ 880 MB,ж„ҹи§үжҳҜеҚіж—¶е“Қеә”;3B жЁЎеһӢзәҰ 2 GB,жҳҺжҳҫжӣҙиҝһиҙҜ;7B жЁЎеһӢи¶…иҝҮ 4 GB,йңҖиҰҒдёҖеқ—еғҸж ·зҡ„жҳҫеҚЎгҖӮиҙЁйҮҸйҡҸдҪ“з§ҜжҸҗеҚҮ,дҪҶзӯүеҫ…ж—¶й—ҙд№ҹйҡҸд№ӢеўһеҠ гҖӮ Qwen 2.5 Coder 1.5B дё“й—Ёй’ҲеҜ№д»Јз Ғи®ӯз»ғ,д»Ҙе…¶дҪ“з§ҜиҖҢиЁҖиЎЁзҺ°еҮәиүІ——еҰӮжһңдҪ йңҖиҰҒд»Јз ҒиЎҘе…ЁгҖҒе°ҸеҮҪж•°е’Ң shell еҚ•иЎҢе‘Ҫд»Ө,е®ғжҳҜжҖ§д»·жҜ”жңҖй«ҳзҡ„йҖүжӢ©гҖӮ Phi 3.5 Mini еңЁйңҖиҰҒжЁЎеһӢзңҹжӯЈ"зҗҶи§Ј"д»Јз ҒиҖҢдёҚеҸӘжҳҜз”ҹжҲҗд»Јз Ғж—¶,жҳҜжӣҙејәзҡ„е…ЁиғҪйҖүжүӢгҖӮдёӨиҖ…йғҪж— жі•еңЁеӨ§еһӢд»Јз Ғеә“дёҠеҸ–д»ЈеүҚжІҝжЁЎеһӢгҖӮ иҝҷдәӣжҳҜеңЁдёҖдёӘжӣҙеӨ§зҡ„жҺЁзҗҶжЁЎеһӢиҫ“еҮәдёҠи®ӯз»ғеҮәжқҘзҡ„е°ҸжЁЎеһӢ,еӣ жӯӨе®ғ们дјҡе…ҲдёҖжӯҘжӯҘжҺЁжј”й—®йўҳ,еҶҚз»ҷеҮәзӯ”жЎҲгҖӮиҝҷи®©е®ғ们еңЁи§Ји°ңе’ҢеӨҡжӯҘйҖ»иҫ‘дёҠзҡ„иЎЁзҺ°дјҳдәҺеҗҢзӯүдҪ“з§Ҝзҡ„жЁЎеһӢ——дҪҶд№ҹжӣҙж…ў,еӣ дёәе®ғ们иҰҒиҠұиҙ№ token"еӨ§еЈ°жҖқиҖғ"гҖӮдҪҝз”Ёе®ғ们时иҜ·и°ғй«ҳ max tokens ,еҗҰеҲҷеӣһзӯ”дјҡиў«дёӯйҖ”жҲӘж–ӯгҖӮ Qwen 2.5 зі»еҲ—еңЁиҝҷж–№йқўжҳҜжңҖејәзҡ„еӨҡиҜӯиЁҖйҖүйЎ№,1.5B е’Ң 3B зүҲжң¬иҰҶзӣ–дәҶе№ҝжіӣзҡ„иҜӯиЁҖгҖӮ Phi 3.5 Mini д№ҹиғҪеҫҲеҘҪең°еӨ„зҗҶ 20 еӨҡз§ҚиҜӯиЁҖгҖӮе°ҸжЁЎеһӢеңЁйқһиӢұиҜӯеңәжҷҜдёӯжӣҙе®№жҳ“"и·‘еҒҸ",жүҖд»ҘжҸҗзӨәиҜҚиҰҒеҶҷеҫ—жҳҺзЎ®——жҠҠзі»з»ҹжҸҗзӨәиҜҚи®ҫдёә“е§Ӣз»Ҳз”Ёдёӯж–Үеӣһзӯ””жҜ”еңЁеҜ№иҜқдёӯйҖ”иҰҒжұӮж•ҲжһңжӣҙеҘҪгҖӮ еҸӘжңүжӢҘжңүеӨ§зәҰ 5 GB жҲ–жӣҙеӨҡеҸҜз”Ёжҳҫеӯҳ зҡ„зӢ¬з«ӢжҳҫеҚЎжүҚиЎҢ——д»…жқғйҮҚжң¬иә«е°ұзәҰ 4.3 GB,иҝҳдёҚз®—еҜ№иҜқзј“еӯҳгҖӮMistral 7B е’Ң DeepSeek R1 7B и’ёйҰҸзүҲйғҪеңЁйҖүжӢ©еҷЁдёӯ,дҫӣиғҪжүҝеҸ—зҡ„и®ҫеӨҮдҪҝз”ЁгҖӮеңЁйӣҶжҲҗжҳҫеҚЎдёҠе®ғ们иҰҒд№ҲеҲҶй…ҚеӨұиҙҘ,иҰҒд№ҲиҝҗиЎҢжһҒж…ў,иҝҷз§Қжғ…еҶөдёӢ 3B жЁЎеһӢжҳҜжӣҙеҘҪзҡ„жҠҳдёӯж–№жЎҲгҖӮ дҪ“з§Ҝе°Ҹеҫ—еӨҡзҡ„жЁЎеһӢ,еӣ дёәе®ғиҝҗиЎҢеңЁ CPU дёҠ:SmolLM2 135M(зәҰ 70 MB)гҖҒSmolLM2 360M(зәҰ 200 MB,жҺЁиҚҗй»ҳи®Ө)гҖҒQwen 1.5 0.5B,д»ҘеҸҠеҶ…еӯҳе……иЈ•и®ҫеӨҮеҸҜз”Ёзҡ„ SmolLM2 1.7BгҖӮе®ғ们жҳҺжҳҫжҜ”жЎҢйқўз«ҜжЁЎеһӢз®ҖеҚ•,дҪҶеҸҜд»ҘеңЁд»»дҪ•и®ҫеӨҮдёҠиҝҗиЎҢ——ж— йңҖ WebGPUгҖӮ жҜҸдёӘжЁЎеһӢеҲҶеҲ«зј“еӯҳ,жүҖд»Ҙ第дёҖж¬ЎеҲҮжҚўеҲ°ж–°жЁЎеһӢдјҡдёӢиҪҪ,д№ӢеҗҺжҜҸж¬ЎеҲҮжҚўйғҪжҳҜеҚіж—¶зҡ„гҖӮе·Ізј“еӯҳзҡ„жЁЎеһӢдјҡеңЁйҖүжӢ©еҷЁдёӯж Үи®°еҮәжқҘ,ж–№дҫҝдҪ еңЁеҝ«йҖҹжЁЎеһӢе’ҢжӣҙејәжЁЎеһӢд№Ӣй—ҙиҮӘеҰӮеҲҮжҚў,иҖҢдёҚеҝ…йҮҚеӨҚдёӢиҪҪгҖӮ е®ғжҳҜжЁЎеһӢиғҪеҗҢж—¶"и®°дҪҸ"еӨҡе°‘ж–Үжң¬——дҪ зҡ„зі»з»ҹжҸҗзӨәиҜҚгҖҒзӣ®еүҚзҡ„еҜ№иҜқеҶ…е®№,д»ҘеҸҠжЁЎеһӢиҮӘе·ұзҡ„еӣһеӨҚйғҪз®—еңЁеҶ…гҖӮ Phi 3.5 Mini еңЁиҝҷйҮҢз©әй—ҙжңҖеӨ§,иҫҫеҲ° 128K tokens;иҖҢ Phi 3 Mini 4K зӯүиҫғж—§зҡ„жЁЎеһӢеҸӘжңүеҮ еҚғ tokensгҖӮеҪ“й•ҝеҜ№иҜқејҖе§Ӣ"и·‘йўҳ"ж—¶,йҮҚзҪ®еҜ№иҜқ,еҸӘзІҳиҙҙзңҹжӯЈйҮҚиҰҒзҡ„еҶ…е®№гҖӮ иҝҷжҳҜйҮҸеҢ–ж–№ејҸ:жқғйҮҚиў«еҺӢзј©еҲ° 4 дҪҚ,и®Ўз®—ж—¶дҪҝз”Ё 16 дҪҚжө®зӮ№ж•°гҖӮиҝҷжӯЈжҳҜи®©дёҖдёӘж•°еҚҒдәҝеҸӮж•°зҡ„жЁЎеһӢе°ҸеҲ°иғҪдёӢиҪҪе№¶иЈ…е…ҘжөҸи§ҲеҷЁеҶ…еӯҳзҡ„еҺҹеӣ гҖӮд»Јд»·жҳҜе°‘йҮҸзІҫеәҰжҚҹеӨұ,дҪҶзӣёжҜ”е®Ңе…Ёж— жі•иҝҗиЎҢжЁЎеһӢиҖҢиЁҖ,иҝҷжҳҜеҫҲеҲ’з®—зҡ„жқғиЎЎгҖӮ 03. йҡҗз§ҒдёҺж•°жҚ® дёҚдјҡгҖӮ жҸҗзӨәиҜҚе’ҢеӣһеӨҚж°ёиҝңдёҚдјҡзҰ»ејҖиҝҷдёӘж ҮзӯҫйЎөгҖӮиҝҷдёӘе·Ҙе…·еҸ‘еҮәзҡ„е”ҜдёҖзҪ‘з»ңиҜ·жұӮ,жҳҜд»Һ Hugging Face CDN дёҖж¬ЎжҖ§дёӢиҪҪжЁЎеһӢжқғйҮҚ;д№ӢеҗҺдҪ еҸҜд»Ҙе®Ңе…Ёж–ӯзҪ‘,е®ғдҫқз„¶иғҪжӯЈеёёе·ҘдҪңгҖӮиҝҷйҮҢжІЎжңүжҺЁзҗҶжңҚеҠЎеҷЁ,еӣ дёәдҪ зҡ„и®ҫеӨҮжң¬иә«е°ұжҳҜжҺЁзҗҶжңҚеҠЎеҷЁгҖӮ еҜ№иҜқеҶ…е®№еӯҳеңЁдәҺйЎөйқўдёӯ,еҪ“дҪ е…ій—ӯжҲ–йҮҚзҪ®ж ҮзӯҫйЎөж—¶е°ұдјҡж¶ҲеӨұгҖӮжІЎжңүд»»дҪ•еҶ…е®№иў«дёҠдј ,д№ҹжІЎжңүеҶҷе…Ҙд»»дҪ•ж•°жҚ®еә“,жөҸи§ҲеҷЁеӯҳеӮЁдёӯдҝқз•ҷзҡ„еҸӘжңүдҪ зҡ„дё»йўҳйҖүжӢ©,д»ҘеҸҠдёҖд»Ҫ"е“ӘдәӣжЁЎеһӢе·Із»Ҹзј“еӯҳиҝҮ"зҡ„и®°еҪ•——з»қдёҚеҢ…еҗ«ж¶ҲжҒҜеҶ…е®№гҖӮеҰӮжһңжғідҝқз•ҷеҜ№иҜқ,дҪҝз”Ё "еӨҚеҲ¶и®°еҪ•" жҲ– "↓ .md" еҜјеҮәеҚіеҸҜгҖӮ жҳҜзңҹзҡ„з§ҒеҜҶ,иҖҢдё”дҪ еҸҜд»ҘиҮӘе·ұйӘҢиҜҒ,дёҚеҝ…еҸӘеҗ¬еҲ«дәәзҡ„иҜҙжі•гҖӮйЎөйқўдёӯжІЎжңүд»»дҪ•еҸҜз”ЁжқҘеҸ‘йҖҒж•°жҚ®зҡ„ API key,д№ҹжІЎжңүеҸҜд»ҘеҸ‘йҖҒж•°жҚ®зҡ„зӣ®ж Үең°еқҖ,иҖҢдё”ж–ӯзҪ‘еҗҺжЁЎеһӢдҫқз„¶иғҪеӣһзӯ”——иҝҷеҜ№дә‘з«ҜиҒҠеӨ©жңәеҷЁдәәжқҘиҜҙжҳҜдёҚеҸҜиғҪеҒҡеҲ°зҡ„гҖӮ дёӨз§Қж–№жі•гҖӮжү“ејҖжөҸи§ҲеҷЁзҡ„ејҖеҸ‘иҖ…е·Ҙе…·,еҲҮеҲ° Network ж ҮзӯҫйЎө,еҸ‘йҖҒж¶ҲжҒҜж—¶и§ӮеҜҹе®ғ:жЁЎеһӢеҠ иҪҪе®ҢжҲҗеҗҺ,дёҚдјҡеҮәзҺ°д»»дҪ•зҪ‘з»ңиҜ·жұӮгҖӮжҲ–иҖ…е№Іи„Ҷе…ій—ӯ Wi-Fi з»§з»ӯиҒҠеӨ©——еҰӮжһңдҫқз„¶иғҪеҫ—еҲ°еӣһзӯ”,иҜҙжҳҺе®ғжҳҜеңЁжң¬ең°з”ҹжҲҗзҡ„гҖӮ д»ҺжҠҖжңҜдёҠи®І,д»»дҪ•еҶ…е®№йғҪдёҚдјҡзҰ»ејҖдҪ зҡ„и®ҫеӨҮ,еӣ жӯӨдёҚеӯҳеңЁеҗ‘第дёүж–№дј иҫ“зҡ„й—®йўҳ,д№ҹдёҚйңҖиҰҒдёҺд»»дҪ•еӨ„зҗҶиҖ…зӯҫзҪІеҚҸи®®——иҝҷжӯЈжҳҜжң¬ең°жЁЎеһӢеҜ№ж•Ҹж„ҹж–Үжң¬еҫҲжңүеҗёеј•еҠӣзҡ„еҺҹеӣ гҖӮиҜқиҷҪеҰӮжӯӨ,дҪ жүҖеңЁз»„з»Үзҡ„ж”ҝзӯ–д»Қз„¶йҖӮз”Ё,жүҖд»ҘеңЁиҝҷйҮҢзІҳиҙҙе®ўжҲ·ж•°жҚ®д№ӢеүҚ,иҜ·е…ҲзЎ®и®ӨдёҖдёӢзӣёе…іи§„е®ҡгҖӮ дёӢиҪҪжқғйҮҚеҸӘжҳҜдёҖж¬Ўжҷ®йҖҡзҡ„ж–Үд»¶иҜ·жұӮ,еҸ‘еҫҖдёҖдёӘ CDN,иҜҘ CDN зңӢеҲ°зҡ„е°ұжҳҜ"жңүдёҖдёӘж–Үд»¶иў«иҺ·еҸ–",е’ҢиҺ·еҸ–д»»дҪ•еӣҫзүҮжҲ–и„ҡжң¬жІЎжңүеҢәеҲ«гҖӮе®ғдёҚжҗәеёҰд»»дҪ•жҸҗзӨәиҜҚгҖҒеҜ№иҜқеҶ…е®№,д№ҹжІЎжңүи¶…еҮәжҷ®йҖҡзҪ‘з»ңиҜ·жұӮзҡ„иә«д»ҪдҝЎжҒҜ——иҖҢдё”жҜҸдёӘжЁЎеһӢеҸӘеҸ‘з”ҹдёҖж¬Ў,дёҚжҳҜжҜҸжқЎж¶ҲжҒҜйғҪеҸ‘з”ҹгҖӮ еҸҜд»Ҙ,дҪҶйҡҗз§ҒзӘ—еҸЈеңЁе…ій—ӯж—¶дјҡжё…йҷӨеӯҳеӮЁ,жүҖд»ҘдёӢж¬ЎдҪҝз”Ёж—¶жЁЎеһӢдјҡйҮҚж–°дёӢиҪҪгҖӮеҰӮжһңдҪ жү“з®—з»ҸеёёдҪҝз”Ё,жҷ®йҖҡзӘ—еҸЈиғҪдҝқз•ҷзј“еӯҳ,зңҒеҺ»зӯүеҫ…ж—¶й—ҙгҖӮ 04. йҖҹеәҰдёҺзЎ¬д»¶ еҜ№дәҺжЎҢйқўж ҮзӯҫйЎө,йңҖиҰҒдёҖеқ—ж”ҜжҢҒ WebGPUгҖҒе№¶жңүи¶іеӨҹз©әй—ІеҶ…еӯҳиҝҗиЎҢжЁЎеһӢзҡ„ GPU:1B жЁЎеһӢеӨ§зәҰйңҖиҰҒ 1 GB,3B жЁЎеһӢйңҖиҰҒ 2–3 GB,7B жЁЎеһӢеҲҷйңҖиҰҒ 5 GB д»ҘдёҠгҖӮйӣҶжҲҗжҳҫеҚЎеҸҜд»ҘиҪ»жқҫеә”еҜ№ 0.5B–1B жЁЎеһӢгҖӮеҰӮжһңдҪ зҡ„и®ҫеӨҮйғҪдёҚж»Ўи¶і,Mobile / WASM ж ҮзӯҫйЎөеҮ д№ҺеҸҜд»ҘеңЁд»»дҪ•и®ҫеӨҮзҡ„ CPU дёҠиҝҗиЎҢгҖӮ еңЁзҺ°д»ЈзӢ¬з«ӢжҳҫеҚЎдёҠ,1B жЁЎеһӢзҡ„з”ҹжҲҗйҖҹеәҰйҖҡеёёжҜ”дҪ йҳ…иҜ»зҡ„йҖҹеәҰиҝҳеҝ«,3B жЁЎеһӢд№ҹдҫқз„¶ж„ҹи§үеғҸжӯЈеёёеҜ№иҜқгҖӮеңЁйӣҶжҲҗжҳҫеҚЎдёҠ,йў„жңҹжҳҜеҸҜиҜ»дҪҶдёҚз®—еҝ«,иҖҢеңЁ CPU / WASM жЁЎејҸдёӢ,йў„жңҹжҳҜжҜҸз§’еҮ дёӘиҜҚгҖӮиҒҠеӨ©жЎҶдёӢж–№зҡ„е®һж—¶ tok/s и®Ўж•°еҷЁдјҡеҮҶзЎ®жҳҫзӨәдҪ и®ҫеӨҮзҡ„иЎЁзҺ°гҖӮ йҖҡеёёжҳҜжЁЎеһӢеҜ№зЎ¬д»¶жқҘиҜҙеӨӘеӨ§дәҶ,еҶ…еӯҳеңЁиў«еҸҚеӨҚи°ғеәҰиҖҢдёҚжҳҜиў«жңүж•ҲеҲ©з”ЁгҖӮйҷҚдҪҺдёҖдёӘдҪ“з§ҜжЎЈж¬ЎгҖҒе…ій—ӯе…¶д»–еҚ з”Ё GPU зҡ„ж ҮзӯҫйЎө——и§Ҷйў‘йҖҡиҜқе’Ң 3D йЎөйқўдјҡдәүжҠўеҗҢж ·зҡ„еҶ…еӯҳ——е№¶зЎ®и®Ө笔记жң¬жІЎжңүеӨ„дәҺйҷҗеҲ¶ GPU жҖ§иғҪзҡ„зңҒз”өжЁЎејҸгҖӮй•ҝеҜ№иҜқд№ҹдјҡжӢ–ж…ўйҖҹеәҰ,еӣ дёәж•ҙдёӘеҺҶеҸІи®°еҪ•йғҪиҰҒиў«йҮҚж–°еӨ„зҗҶ;йҮҚзҪ®еҜ№иҜқеҚіеҸҜжҒўеӨҚйҖҹеәҰгҖӮ е®ғи§ҰеҸҠдәҶ max tokens дёҠйҷҗ,й»ҳи®ӨеҖјжҳҜ 1024гҖӮеҜ№дәҺи„ҡжң¬гҖҒй•ҝд»Јз ҒжҲ–иҜҰз»Ҷж–ҮжЎЈ,иҜ·еңЁ Send жҢүй’®ж—Ғе°Ҷе…¶и°ғй«ҳеҲ° 2048гҖҒ4096 жҲ– 8192гҖӮиҝҷжҳҜдёҖдёӘеҲ»ж„Ҹи®ҫзҪ®зҡ„дёҠйҷҗ,иҖҢдёҚжҳҜж•…йҡң——жІЎжңүе®ғ,е°ҸжЁЎеһӢеҸҜиғҪдјҡй•ҝзҜҮеӨ§и®әгҖҒжІЎе®ҢжІЎдәҶгҖӮ е®ғжҺ§еҲ¶йҮҮж ·зҡ„"еӨ§иғҶзЁӢеәҰ"гҖӮиҫғдҪҺзҡ„еҖј(зәҰ 0.2)дјҡи®©жЁЎеһӢйҮҚеӨҚдҪҶеҸҜйў„жөӢ,иҝҷйҖӮеҗҲз”ЁдәҺжҸҗеҸ–дҝЎжҒҜгҖҒJSON е’Ңд»Јз ҒгҖӮиҫғй«ҳзҡ„еҖј(0.8 еҸҠд»ҘдёҠ)дјҡдә§з”ҹжӣҙеӨҡж ·еҢ–зҡ„ж–Үеӯ—,дҪҶеҮҶзЎ®жҖ§дјҡдёӢйҷҚгҖӮй»ҳи®ӨеҖј 0.7 жҳҜиҒҠеӨ©еңәжҷҜжҜ”иҫғеҗҲзҗҶзҡ„жҠҳдёӯгҖӮ йҖҡеёёеҝ«еҘҪеҮ еҖҚ,жңүж—¶з”ҡиҮіиғҪеҝ«дёҖдёӘж•°йҮҸзә§,иҝҷд№ҹжҳҜдёәд»Җд№Ҳ WebGPU ж ҮзӯҫйЎөжҳҜжЎҢйқўз«Ҝзҡ„й»ҳи®ӨйҖүйЎ№гҖӮWASM жЁЎејҸзҡ„еӯҳеңЁ,жҳҜдёәдәҶи®©е·Ҙе…·еңЁжІЎжңү WebGPU зҡ„зЎ¬д»¶е’ҢжөҸи§ҲеҷЁдёҠд№ҹиғҪиҝҗиЎҢ——е®ғд»ҘйҖҹеәҰжҚўеҸ–дәҶеҮ д№Һж— еӨ„дёҚеңЁзҡ„е…је®№жҖ§гҖӮ жЁЎеһӢж”ҫдёҚдёӢдәҶгҖӮйҖүжӢ©дёҖдёӘжӣҙе°Ҹзҡ„жЁЎеһӢ,е…ій—ӯе…¶д»–еҚ з”Ё GPU зҡ„ж ҮзӯҫйЎө,еҶҚиҜ•дёҖж¬Ў——й”ҷиҜҜдҝЎжҒҜдјҡжҢҮеҮәе…·дҪ“и§ҰеҸҠдәҶе“ӘдёӘйҷҗеҲ¶гҖӮеңЁдјҡеңЁйӣҶжҲҗжҳҫеҚЎе’ҢзӢ¬з«ӢжҳҫеҚЎд№Ӣй—ҙеҲҮжҚўзҡ„笔记жң¬дёҠ,еңЁзі»з»ҹи®ҫзҪ®дёӯејәеҲ¶жөҸи§ҲеҷЁдҪҝз”ЁзӢ¬з«ӢжҳҫеҚЎ,еҫҖеҫҖиғҪеҪ»еә•и§ЈеҶій—®йўҳгҖӮ еңЁз”ҹжҲҗиҝҮзЁӢдёӯ,дјҡзҡ„——дҪ жҳҜеңЁиҮӘе·ұзҡ„иҠҜзүҮдёҠиҝҗиЎҢдёҖдёӘзҘһз»ҸзҪ‘з»ң,йЈҺжүҮеҸҜиғҪдјҡеғҸзҺ©жёёжҲҸж—¶дёҖж ·иҪ¬иө·жқҘгҖӮе®ғеҸӘеңЁз”ҹжҲҗеӣһеӨҚж—¶иҖ—з”ө;жЁЎеһӢй—ІзҪ®еңЁеҶ…еӯҳдёӯж—¶дёҚдјҡдә§з”ҹд»»дҪ•йўқеӨ–еҠҹиҖ—гҖӮ 05. жөҸи§ҲеҷЁгҖҒеӯҳеӮЁдёҺж•…йҡңжҺ’жҹҘ Chrome е’Ң Edge иҮӘ 113 зүҲжң¬иө·,еңЁ WindowsгҖҒmacOS е’Ң ChromeOS дёҠй»ҳи®Өе·Іж”ҜжҢҒ,Android 12+ дёҠеҲҷд»Һ Chrome 121 иө·ж”ҜжҢҒгҖӮSafari еңЁ macOS Tahoe 26гҖҒiOS 26 е’Ң iPadOS 26 дёӯй»ҳи®ӨеҗҜз”ЁгҖӮFirefox еңЁ Windows дёҠд»Һ 141 зүҲжң¬иө·ж”ҜжҢҒ,еңЁ Apple Silicon Mac дёҠд»Һ 147 зүҲжң¬иө·ж”ҜжҢҒ,Linux е’Ң Android дёҠд»ҚеӨ„дәҺ Nightly йҳ¶ж®өгҖӮдёҚеңЁжӯӨеҲ—иЎЁдёӯзҡ„жөҸи§ҲеҷЁд»ҚеҸҜдҪҝз”Ё Mobile / WASM ж ҮзӯҫйЎөгҖӮ дҪ зҡ„жөҸи§ҲеҷЁ,жҲ–иҖ…иҜҘи®ҫеӨҮзҡ„жҳҫеҚЎй©ұеҠЁ,жІЎжңүејҖж”ҫ WebGPU ж”ҜжҢҒ——иҝҷеңЁиҫғж—§зҡ„жөҸи§ҲеҷЁзүҲжң¬гҖҒйғЁеҲҶй©ұеҠЁзҡ„ Linux зі»з»ҹгҖҒжҹҗдәӣиҷҡжӢҹжңә,д»ҘеҸҠзҰҒз”ЁдәҶиҜҘеҠҹиғҪзҡ„дјҒдёҡзӯ–з•ҘзҺҜеўғдёӯеҫҲеёёи§ҒгҖӮйЎөйқўжЈҖжөӢеҲ°иҝҷдёҖзӮ№еҗҺдјҡиҮӘеҠЁжҠҠдҪ еҲҮжҚўеҲ° Mobile / WASM ж ҮзӯҫйЎө,и®©дҪ дҫқз„¶иғҪдҪҝз”Ёе·Ҙе…·;еҰӮжһңжғіиҰҒ GPU йҖҹеәҰ,еҚҮзә§жөҸи§ҲеҷЁе°ұжҳҜи§ЈеҶіеҠһжі•гҖӮ еҸҜд»Ҙ,йҖҡиҝҮ Mobile / WASM ж ҮзӯҫйЎө,дҪҶжңүдёҖдёӘжіЁж„ҸдәӢйЎ№:iOS Safari еҜ№жҜҸдёӘж ҮзӯҫйЎөејәеҲ¶ж–ҪеҠ зәҰ 1–1.5 GB зҡ„зЎ¬жҖ§еҶ…еӯҳдёҠйҷҗ,жүҖд»ҘеҸӘжңүжңҖе°Ҹзҡ„жЁЎеһӢ(SmolLM2 135M)иғҪеҸҜйқ еҠ иҪҪгҖӮжӣҙеӨ§зҡ„жЁЎеһӢдјҡи®©ж ҮзӯҫйЎөеҙ©жәғ,жҸҗзӨә“a problem repeatedly occurred”гҖӮеҰӮжһңйңҖиҰҒжӣҙејәзҡ„иғҪеҠӣ,иҜ·еңЁжЎҢйқўз«Ҝжү“ејҖжӯӨйЎөйқўгҖӮ еҸҜд»ҘгҖӮAndroid 12 еҸҠжӣҙй«ҳзүҲжң¬дёҠзҡ„ Chrome еңЁеӨ§еӨҡж•°иҝ‘жңҹиҠҜзүҮдёҠж”ҜжҢҒ WebGPU,дёҚж”ҜжҢҒж—¶еҲҷз”ұ WASM ж ҮзӯҫйЎөеңЁ CPU дёҠиҝҗиЎҢгҖӮиҜ·дјҳе…ҲйҖүжӢ©е°ҸжЁЎеһӢ——жүӢжңәзҡ„еҶ…еӯҳйҷҗеҲ¶еҫҖеҫҖжҜ”жЁЎеһӢиҙЁйҮҸжӣҙж—©жҲҗдёәз“¶йўҲгҖӮ дјҒдёҡд»ЈзҗҶе’ҢеҶ…е®№иҝҮж»ӨеҷЁз»ҸеёёдјҡеұҸи”Ҫ Hugging Face CDN(жқғйҮҚжӯЈжҳҜд»ҺиҝҷйҮҢдёӢиҪҪзҡ„),жңүдәӣиҝҳдјҡжЈҖжҹҘе№¶дёӯж–ӯеӨ§ж–Үд»¶дёӢиҪҪгҖӮиҝҷжҳҜзҪ‘з»ңзӯ–з•Ҙй—®йўҳ,иҖҢдёҚжҳҜжөҸи§ҲеҷЁй—®йўҳ——еҗҢдёҖдёӘйЎөйқўеңЁе®¶еәӯзҪ‘з»ңжҲ–жүӢжңәзғӯзӮ№дёҠйҖҡеёёиғҪз«ӢеҚіжӯЈеёёе·ҘдҪңгҖӮ еӯҳеӮЁеңЁжөҸи§ҲеҷЁй’ҲеҜ№иҜҘзҪ‘з«ҷзҡ„ Cache Storage дёӯ,иҖҢдёҚжҳҜдҪ зҡ„ж–Үд»¶зі»з»ҹйҮҢгҖӮжё…йҷӨзҪ‘з«ҷж•°жҚ®(жҲ–й’ҲеҜ№иҜҘеҹҹеҗҚжё…йҷӨ“Cookie еҸҠе…¶д»–зҪ‘з«ҷж•°жҚ®”)дјҡз§»йҷӨжүҖжңүе·Ізј“еӯҳзҡ„жЁЎеһӢе№¶йҮҠж”ҫз©әй—ҙ;д№ӢеҗҺе·Ҙе…·зҡ„иЎЁзҺ°е°ұеҰӮеҗҢйҰ–ж¬Ўи®ҝй—®дёҖж ·гҖӮжөҸи§ҲеҷЁй…ҚзҪ®ж–Үд»¶д№ӢеӨ–зҡ„д»»дҪ•еҶ…е®№йғҪдёҚдјҡеҸ—еҲ°еҪұе“ҚгҖӮ дјҡзҡ„гҖӮеҪ“зЈҒзӣҳз©әй—ҙдёҚи¶іжҲ–й…ҚзҪ®ж–Үд»¶иў«жё…зҗҶж—¶,жөҸи§ҲеҷЁдјҡжё…йҷӨзј“еӯҳж•°жҚ®,йҡҗз§ҒзӘ—еҸЈеңЁе…ій—ӯж—¶д№ҹдјҡдёўејғзј“еӯҳгҖӮеҰӮжһңдёҖдёӘжҳЁеӨ©иҝҳиғҪзһ¬й—ҙеҠ иҪҪзҡ„жЁЎеһӢзӘҒз„¶еҸҲејҖе§ӢйҮҚж–°дёӢиҪҪ,еҺҹеӣ е°ұеңЁиҝҷйҮҢ——е№¶дёҚжҳҜеҮәдәҶд»Җд№Ҳж•…йҡңгҖӮ еҸӘдјҡеҚ з”ЁдҪ е®һйҷ…еҠ иҪҪиҝҮзҡ„жЁЎеһӢжүҖйңҖзҡ„з©әй—ҙ:жңҖе°Ҹзҡ„зәҰ 70 MB,7B жЁЎеһӢзәҰ 4.5 GBгҖӮдҪ иҜ•з”ЁиҝҮзҡ„жҜҸдёӘжЁЎеһӢйғҪдјҡдёҖзӣҙдҝқз•ҷ,зӣҙеҲ°дҪ жё…йҷӨзҪ‘з«ҷж•°жҚ®,жүҖд»Ҙе°қиҜ•еӨҡдёӘжЁЎеһӢзЎ®е®һдјҡзҙҜз§ҜеҚ з”Ё——еңЁзЈҒзӣҳз©әй—ҙзҙ§еј зҡ„и®ҫеӨҮдёҠеҖјеҫ—з•ҷж„ҸгҖӮ 06. е®ғж“…й•ҝд»Җд№Ҳ——д»ҘеҸҠдёҚж“…й•ҝд»Җд№Ҳ еӣ дёәе®ғзҡ„дҪ“з§Ҝе°ҸдәҶдёүдёӘж•°йҮҸзә§гҖӮеүҚжІҝжЁЎеһӢиҝҗиЎҢеңЁжӢҘжңүж•°еҚғдәҝеҸӮж•°зҡ„жңҚеҠЎеҷЁжңәзҫӨдёҠ;жөҸи§ҲеҷЁж ҮзӯҫйЎөйҮҢзҡ„ 1B жЁЎеһӢеҸӘжҳҜе…¶дёӯзҡ„дёҖе°ҸйғЁеҲҶ,иҖҢе®ғиҝҳиғҪжӯЈеёёе·ҘдҪң,е·Із»ҸеҫҲдәҶдёҚиө·дәҶгҖӮжҠҠе®ғеҪ“дҪңдёҖдёӘеӨ„зҗҶе°Ҹд»»еҠЎзҡ„еҝ«йҖҹжң¬ең°еҠ©жүӢ,иҖҢдёҚжҳҜеүҚжІҝжЁЎеһӢзҡ„жӣҝд»Је“ҒгҖӮ зҹӯе°ҸгҖҒиҫ№з•Ңжё…жҷ°зҡ„ж–Үеӯ—е·ҘдҪң:жҠҠдёҖж®өиҜқж”№еҶҷеҫ—жӣҙжӯЈејҸжҲ–жӣҙйҡҸж„ҸгҖҒжҖ»з»“дҪ зІҳиҙҙиҝӣжқҘзҡ„еҶ…е®№гҖҒиө·иҚүеӣһеӨҚгҖҒжҠҠдҝЎжҒҜжҸҗеҸ–дёә JSONгҖҒз®ҖеҚ•ең°и§ЈйҮҠдёҖдёӘжҰӮеҝө,д»ҘеҸҠе°Ҹж®өд»Јз Ғе’ҢжҹҘиҜўиҜӯеҸҘгҖӮиҒҠеӨ©жЎҶдёҠж–№зҡ„йў„и®ҫжҢүй’®жӯЈжҳҜиҝҷдәӣд»»еҠЎзҡ„еҝ«жҚ·ж–№ејҸгҖӮд»»дҪ•йңҖиҰҒжңҖж–°дәӢе®һжҲ–й•ҝй“ҫжқЎжҺЁзҗҶзҡ„д»»еҠЎ,йғҪдёҚйҖӮеҗҲдәӨз»ҷе®ғгҖӮ дёҚиғҪгҖӮе®ғжІЎжңүдә’иҒ”зҪ‘и®ҝй—®жқғйҷҗ,д№ҹж— жі•и®ҝй—®дҪ зҡ„ж–Үд»¶——е®ғеҸӘиғҪзңӢеҲ°дҪ иҫ“е…ҘеҲ°жЎҶйҮҢзҡ„еҶ…е®№гҖӮиҝҷд№ҹжӯЈжҳҜе®ғз§ҒеҜҶзҡ„еҺҹеӣ :еҸҢеҗ‘йғҪжІЎжңүд»»дҪ•йҖҡйҒ“гҖӮеҰӮжһңдҪ жғіи®©е®ғеӨ„зҗҶжҹҗд»Ҫж–ҮжЎЈ,жҠҠзӣёе…ійғЁеҲҶзІҳиҙҙиҝӣеҺ»еҚіеҸҜгҖӮ жҜ”еӨ§жЁЎеһӢжӣҙе®№жҳ“,жҳҜзҡ„гҖӮе°ҸжЁЎеһӢеҜ№еҗҚз§°гҖҒж—ҘжңҹгҖҒж•°еӯ—е’Ңд»»дҪ•дё“дёҡеҶ…е®№йғҪиЎЁзҺ°еҫ—"еҫҲиҮӘдҝЎеҚҙеёёеёёеҮәй”ҷ",иҖҢдё”е®ғж— жі•жҹҘиҜҒд»»дҪ•дҝЎжҒҜгҖӮжҠҠдәӢе®һжҖ§зҡ„з»“и®әеҪ“дҪңйңҖиҰҒж ёе®һзҡ„иҚүзЁҝ,е№¶дјҳе…ҲдҪҝз”ЁйӮЈдәӣеҺҹе§Ӣжқҗж–ҷе°ұеңЁдҪ жҸҗзӨәиҜҚйҮҢзҡ„д»»еҠЎ——ж”№еҶҷгҖҒжҖ»з»“гҖҒжҸҗеҸ–гҖӮ еҸҜд»ҘгҖӮеңЁжЁЎеһӢйҖүжӢ©еҷЁдёӢж–№жү“ејҖ "зі»з»ҹжҸҗзӨәиҜҚ(и§’иүІи®ҫе®ҡ)" ,жҸҸиҝ°е®ғеә”иҜҘеҰӮдҪ•еӣһеә”——дҫӢеҰӮ“дҪ жҳҜдёҖдҪҚиЁҖз®Җж„Ҹиө…зҡ„иө„ж·ұе·ҘзЁӢеёҲ”гҖҒ“е§Ӣз»Ҳз”Ёдёӯж–Үеӣһзӯ””гҖҒ“еҸӘз”Ё JSON еӣһеӨҚ”гҖӮе®ғдјҡеә”з”ЁеҲ°еҜ№иҜқдёӯзҡ„жҜҸдёҖжқЎж¶ҲжҒҜ,жҳҜжҸҗеҚҮе°ҸжЁЎеһӢиҫ“еҮәиҙЁйҮҸжңҖжңүж•Ҳзҡ„ж–№жі•гҖӮ еҜ№дәҺе°ҸеһӢгҖҒзӢ¬з«Ӣзҡ„д»»еҠЎ,еҸҜд»Ҙ——жҜ”еҰӮжӯЈеҲҷиЎЁиҫҫејҸгҖҒshell еҚ•иЎҢе‘Ҫд»ӨгҖҒдёҖдёӘеҮҪж•°,жҲ–иҖ…йў„и®ҫйҮҢзҡ„ KQL жҹҘиҜўгҖӮдҪҝз”Ё Qwen 2.5 Coder жҲ– Phi 3.5 Mini,жҠҠйңҖжұӮеҶҷеҫ—е…·дҪ“дёҖдәӣ,и°ғй«ҳ max tokens д»Ҙе…Қиў«жҲӘж–ӯ,е№¶еңЁиҝҗиЎҢд№ӢеүҚе…ҲжЈҖжҹҘе®ғеҶҷеҮәзҡ„еҶ…е®№гҖӮе®ғе№¶дёҚдәҶи§ЈдҪ д»Јз Ғеә“йҮҢзҡ„д»»дҪ•еҶ…е®№гҖӮ дёҖиҲ¬жқҘиҜҙеҸҜд»Ҙ,дҪҶи®ёеҸҜиҜҒеұһдәҺжЁЎеһӢжң¬иә«,иҖҢдёҚжҳҜиҝҷдёӘйЎөйқў,иҖҢдё”еҗ„дёҚзӣёеҗҢ:Llama йҒөеҫӘ Meta зҡ„зӨҫеҢәи®ёеҸҜиҜҒ,Gemma йҒөеҫӘ Google зҡ„дҪҝз”ЁжқЎж¬ҫ,еӨ§еӨҡж•° Qwen е’Ң SmolLM2 зүҲжң¬йҮҮз”Ё Apache 2.0,Phi йҮҮз”Ё MIT и®ёеҸҜиҜҒгҖӮеҰӮжһңиҫ“еҮәиҰҒз”ЁдәҺдә§е“Ғ,иҜ·жҹҘйҳ…дҪ жүҖдҪҝз”Ёзҡ„е…·дҪ“жЁЎеһӢзҡ„и®ёеҸҜиҜҒгҖӮ 07. жҠҖжңҜеҺҹзҗҶ жЎҢйқўж ҮзӯҫйЎөдҪҝз”ЁжқҘиҮӘ MLC AI зҡ„ WebLLM ,е®ғжҠҠејҖж”ҫжқғйҮҚжЁЎеһӢзј–иҜ‘дёәеҸҜеңЁ WebGPU дёҠиҝҗиЎҢзҡ„еҪўејҸгҖӮз§»еҠЁз«Ҝж ҮзӯҫйЎөдҪҝз”Ё Hugging Face зҡ„ transformers.js,еә•еұӮеҹәдәҺ ONNX Runtime Web,йҖҡиҝҮ WebAssembly еңЁ CPU дёҠжү§иЎҢгҖӮеӣҙз»•е®ғ们зҡ„дёҖеҲҮйғҪеҸӘжҳҜдёҖдёӘеҚ•дёҖ HTML ж–Үд»¶йҮҢзҡ„жҷ®йҖҡ JavaScript е’Ң CSSгҖӮ жқҘиҮӘ Hugging Face CDN,WebGPU ж ҮзӯҫйЎөдҪҝз”Ёйў„зј–иҜ‘зҡ„ MLC ж јејҸ,WASM ж ҮзӯҫйЎөдҪҝз”Ё ONNX ж јејҸгҖӮе®ғ们йғҪжҳҜз”ұ MetaгҖҒAlibabaгҖҒMicrosoftгҖҒGoogle DeepMindгҖҒHugging Face е’Ң DeepSeek еҸ‘еёғзҡ„ејҖж”ҫжқғйҮҚжЁЎеһӢ——д»»дҪ•дәәйғҪеҸҜд»ҘдёӢиҪҪе№¶еңЁжң¬ең°иҝҗиЎҢзҡ„еҗҢдёҖжү№ж–Үд»¶гҖӮ зҗҶеҝөзӣёеҗҢ,дәӨд»ҳж–№ејҸдёҚеҗҢгҖӮйӮЈдәӣжҳҜйңҖиҰҒе®үиЈ…зҡ„жЎҢйқўеә”з”ЁзЁӢеәҸ,з”ұдәҺзӣҙжҺҘдҪҝз”ЁдҪ зҡ„зЎ¬д»¶,е®ғ们е§Ӣз»Ҳдјҡжӣҙеҝ«,д№ҹиғҪж”ҜжҢҒжӣҙеӨ§зҡ„жЁЎеһӢгҖӮиҖҢиҝҷдёӘе·Ҙе…·еңЁжөҸи§ҲеҷЁж ҮзӯҫйЎөйҮҢиҝҗиЎҢжЁЎеһӢ,е®Ңе…ЁдёҚйңҖиҰҒе®үиЈ…д»»дҪ•дёңиҘҝ,иҝҷи®©е®ғжҲҗдёәе°қиҜ•жң¬ең°жЁЎеһӢжӣҙз®ҖеҚ•зҡ„ж–№ејҸ——ж— и®әжҳҜеңЁеҸ—йҷҗи®ҫеӨҮдёҠгҖҒеҲ«дәәзҡ„з”өи„‘дёҠ,иҝҳжҳҜеҚ•зәҜжғізңӢзңӢиҝҷжҳҜжҖҺд№ҲеӣһдәӢгҖӮ еҪ“д»»еҠЎйңҖиҰҒе№ҝеәҰгҖҒдәӢе®һеҮҶзЎ®жҖ§гҖҒй•ҝж–ҮжЎЈгҖҒжңҖж–°дҝЎжҒҜжҲ–дёҘиӮғзҡ„жҺЁзҗҶж—¶——еүҚжІҝжЁЎеһӢдёҖж¬Ўе°ұиғҪе®ҢжҲҗ 1B жЁЎеһӢе®Ңе…ЁеҒҡдёҚеҲ°зҡ„дәӢгҖӮиҖҢеңЁж–Үжң¬ж•Ҹж„ҹгҖҒдҪ еӨ„дәҺзҰ»зәҝзҠ¶жҖҒгҖҒдҪ жғідёҚејҖиҙҰжҲ·е°ұиҺ·еҫ—еҚіж—¶еӣһзӯ”,жҲ–иҖ…д»»еҠЎи¶іеӨҹе°ҸгҖҒе°ҸжЁЎеһӢжң¬иә«е°ұжҳҜжӣҙеҝ«жҚ·зҡ„е·Ҙе…·ж—¶,иҜ·дҪҝз”ЁиҝҷдёӘж–№жЎҲгҖӮ
📥 下载地址(文章结尾)
装机神器,在线重装利器,在线安装一切系统。