Library for conversion between Traditional and Simplified Chinese
9,976
stars
1,794
commits
C++
primary language
Sep 9, 2026
updated

Open Chinese Convert (OpenCC, 開放中文轉換) is an open source project for high-quality conversion between Traditional Chinese, Simplified Chinese, Japanese Shinjitai, and regional wording across Mainland China, Taiwan, and Hong Kong. It provides dictionaries, a reusable library, conversion tools, and dictionary generation tools.
Open Chinese Convert(OpenCC,開放中文轉換) 是一個開源中文轉換項目,支持繁體中文、簡體中文、日文新字體,以及中國大陸、臺灣、香港等地區習慣用詞之間的高品質轉換,並提供詞典、可重用庫、轉換工具及詞典生成工具。
Discussion (Telegram) 討論區(Telegram): https://t.me/open_chinese_convert
詳情參閱 OpenCC 設計思想及地區詞收錄標準。 For details, see Design Principles and Regional Phrase Criteria.
brew install openccwinget install openccnpm install -g opencc 命令可安裝 OpenCC Node.js CLInpm install -g opencc opencc-jieba 命令可同時安裝 OpenCC Node.js CLI 及 Jieba 分詞插件pip install opencc 命令可安裝 Python API 及 Python CLIOpenCC 1.4.2 大幅加速了轉換熱路徑(純文字語料的整體轉換時間最多降至原本的
1/7),修復了大端序平台載入 legacy .ocd 字典得到空字典的問題,並包含一批詞庫
修正;C++ ABI 與 1.4.1 相同(SOVERSION 1.4),自 1.4.0/1.4.1 升級的下游 C++
程式無需重新連結。
opencc, opencc-jieba, and libopencc* deb packages for one architecture, with a SHA256SUMS file.https://opencc.js.org/converter?config=s2t
npm install opencc
The npm package supports Node.js >=20.17. The native addon is installed
through prebuilt @opencc/opencc-<platform>-<arch> packages covering macOS
(x64/arm64), Linux (x64/arm64), and Windows (x64). On platforms without a
prebuilt binary package, npm install builds from source with Bazel —
compiling the addon and regenerating the dictionaries — which requires a
C++ toolchain and network access; Bazel is found on PATH (bazel or
bazelisk) or fetched automatically via npx, and Bazel downloads its own
hermetic Python toolchain for the dictionary generation.
Bun and Deno can also use the npm package through their npm compatibility support.
To install the npm CLI:
npm install -g opencc
opencc -c s2t.json -i input.txt -o output.txt
The npm CLI supports basic text conversion. Plugins, --inspect,
--segmentation, and --ambiguities require the native OpenCC CLI.
import { OpenCC } from 'opencc';
async function main() {
const converter: OpenCC = new OpenCC('s2t.json');
const result: string = await converter.convertPromise('汉字');
console.log(result); // 漢字
}
Inline configurations can be passed with OpenCC.fromConfig:
import { OpenCC } from 'opencc';
const converter = OpenCC.fromConfig({
name: 'Demo Inline Config',
conversion_chain: [{
dict: { type: 'inline', entries: { '鼠标': '滑鼠' } },
}],
});
console.log(converter.convertSync('鼠标')); // 滑鼠
pip install opencc (Windows, Linux, macOS)
import opencc
converter = opencc.OpenCC('s2t.json')
converter.convert('汉字') # 漢字
The Python package also installs a basic CLI:
pip install opencc
opencc -c s2t.json -i input.txt -o output.txt
The Python CLI supports basic text conversion, --include-tofu-risk-dictionaries,
and --resource-zip. Diagnostic modes such as --inspect, --segmentation,
and --ambiguities still require the native OpenCC CLI.
#include "opencc.h"
int main() {
const opencc::SimpleConverter converter("s2t.json");
converter.Convert("汉字"); // 漢字
return 0;
}
When OpenCC is embedded in a server binary or self-contained application, the JSON config can stay small while dictionary resources are loaded from explicit resource directories:
#include <memory>
#include <vector>
#include "SimpleConverter.hpp"
int main() {
auto resources = std::make_shared<opencc::FilesystemResourceProvider>(
std::vector<std::string>{
"/opt/my-app/opencc",
"/opt/my-app/plugins/opencc-jieba",
"/usr/share/opencc",
});
const opencc::SimpleConverter converter("s2t.json", resources);
converter.Convert("汉字");
return 0;
}
FilesystemResourceProvider searches directories in order. Existing
SimpleConverter("s2t.json") and CLI behavior continue to use the config file
location, current directory, explicit paths, and installed OpenCC data directory
as before.
#include "opencc.h"
int main() {
opencc_t opencc = opencc_open("s2t.json");
const char* input = "汉字";
char* converted = opencc_convert_utf8(opencc, input, strlen(input)); // 漢字
opencc_convert_utf8_free(converted);
opencc_close(opencc);
return 0;
}
Unless otherwise noted, this section describes the native OpenCC CLI built from
the C++ toolchain. The Python and npm CLIs support basic file/stdin conversion
only, plus --include-tofu-risk-dictionaries; the Python CLI also supports
--resource-zip.
opencc --helpopencc_dict --helpOpenCC CLI supports two diagnostic modes that output JSON instead of converted text:
--segmentation — Output segmentation result only (no conversion):
echo "他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题" | opencc -c s2twp.json --segmentation
# {"input":"他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题","segments":["他","只看","了几行","日志",",就","一叶知秋",",猜到","整个","系统","是","数据库","连接池","出了","问题"]}
--inspect — Output full inspection result (segmentation + per-stage conversion + final output):
echo "他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题" | opencc -c s2twp.json --inspect
# {"input":"他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题","segments":["他","只看","了几行","日志",",就","一叶知秋",",猜到","整个","系统","是","数据库","连接池","出了","问题"],"stages":[{"index":1,"segments":["他","只看","了幾行","日誌",",就","一葉知秋",",猜到","整個","系統","是","數據庫","連接池","出了","問題"]},{"index":2,"segments":["他","只看","了幾行","日誌",",就","一葉知秋",",猜到","整個","系統","是","資料庫","連線池","出了","問題"]},{"index":3,"segments":["他","只看","了幾行","日誌",",就","一葉知秋",",猜到","整個","系統","是","資料庫","連線池","出了","問題"]}],"output":"他只看了幾行日誌,就一葉知秋,猜到整個系統是資料庫連線池出了問題"}
# Pretty-print with jq:
echo "他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题" | opencc -c s2twp.json --inspect | jq .
--ambiguities — Convert while marking every output span whose dictionary
match is one-to-many (e.g. 文丑 may be either 文丑 or 文醜), as a stream
of JSONL records:
printf '大战文丑的时候,他的头发很干燥' | opencc -c s2t.json --ambiguities
# {"def":"文丑"}
# {"lit":"大戰"}
# {"amb":{"t":"文丑","s":0}}
# {"lit":"的時候,他的頭髮很乾燥"}
# {"end":{"output_bytes":45,"ambiguities":1,"sources":1}}
Record kinds: {"def"} defines the next source index (each distinct input
word is defined once, before its first reference); {"lit"} is a literal
run of unambiguous output; {"amb":{"t","s"}} is an ambiguous span whose
output text t (the default candidate) came from the source with index s;
the final {"end"} record carries stream totals. Concatenating every lit
and amb.t in order reproduces the plain conversion output exactly.
Streaming is bounded-memory regardless of input size or line length, and
positions are implicit in record order, so consumers can rebuild offsets in
their own string-index units. Single-value conversions (e.g. 头发 →
頭髮) are not flagged.
The record stream is a machine-readable contract: every emitted line is
valid JSON, and the final {"end"} record only appears when the whole
input was processed. Input containing invalid UTF-8 aborts the stream with
an error (unlike plain conversion, which is byte-transparent), and a
missing {"end"} record means the stream is incomplete.
These modes are useful for diagnosing conversion issues:
--segmentation to verify that the input is segmented as expected.--inspect to see which conversion stage produces an unexpected result.--ambiguities to locate one-to-many conversions and resolve each
deduplicated def source to its candidates.Rules:
--segmentation, --inspect, and --ambiguities are mutually exclusive.The following ports are maintained within the OpenCC ecosystem and are generally up to date with current configuration and dictionary data.
These ports are community-maintained and may not always track upstream updates.
s2t.json Simplified Chinese to Traditional Chinese (OpenCC Standard) / 簡體 到 OpenCC 標準繁體t2s.json Traditional Chinese (OpenCC Standard) to Simplified Chinese / OpenCC 標準繁體 到 簡體s2tw.json Simplified Chinese to Traditional Chinese (Taiwan Standard) / 簡體 到 台灣正體tw2s.json Traditional Chinese (Taiwan Standard) to Simplified Chinese / 台灣正體 到 簡體s2hk.json Simplified Chinese to Traditional Chinese (Hong Kong variant) / 簡體 到 香港繁體hk2s.json Traditional Chinese (Hong Kong variant) to Simplified Chinese / 香港繁體 到 簡體s2twp.json Simplified Chinese to Traditional Chinese (Taiwan Standard, with Taiwan Phrases) / 簡體 到 台灣正體(含台灣常用詞彙)tw2sp.json Traditional Chinese (Taiwan Standard) to Simplified Chinese (Mainland China Phrases) / 台灣正體 到 簡體(含中國大陸常用詞彙)t2tw.json Traditional Chinese (OpenCC Standard) to Traditional Chinese (Taiwan Standard) / OpenCC 標準繁體 到 台灣正體tw2t.json Traditional Chinese (Taiwan Standard) to Traditional Chinese (OpenCC Standard) / 台灣正體 到 OpenCC 標準繁體t2hk.json Traditional Chinese (OpenCC Standard) to Traditional Chinese (Hong Kong variant) / OpenCC 標準繁體 到 香港繁體hk2t.json Traditional Chinese (Hong Kong variant) to Traditional Chinese (OpenCC Standard) / 香港繁體 到 OpenCC 標準繁體下列配置文件仍在開發中,歡迎貢獻新詞組:
s2hkp.json Simplified Chinese to Traditional Chinese (Hong Kong variant, with Hong Kong Phrases) / 簡體 到 香港繁體(香港常用詞彙)hk2sp.json Traditional Chinese (Hong Kong variant) to Simplified Chinese (Mainland China Phrases) / 香港繁體 到 簡體(含中國大陸常用詞彙)下列配置文件僅供探索性研究,不建議用於生產環境:
t2jp.json Old Japanese Kanji (Kyūjitai) to New Japanese Kanji (Shinjitai) / 日文舊字體 到 日文新字體jp2t.json New Japanese Kanji (Shinjitai) to Old Japanese Kanji (Kyūjitai) / 日文新字體 到 日文舊字體,並將少量日文詞組轉換爲對應中文通过环境变量OPENCC_DATA_DIR加载指定路径下的配置文件
OPENCC_DATA_DIR=/path/to/your/config/dir opencc --help
配置檔中的字典可使用 type: "inline",直接在 JSON 裡定義小型自訂詞彙,
不必修改外部字典檔。例如在 group.dicts 最前面加入覆寫規則:
{
"conversion_chain": [
{
"dict": {
"type": "group",
"dicts": [
{
"type": "inline",
"entries": {
"麦旋风": "冰炫風",
"服务器": "伺服器"
}
},
{
"type": "ocd2",
"file": "STPhrases.ocd2"
},
{
"type": "ocd2",
"file": "STCharacters.ocd2"
}
]
}
}
]
}
規則與限制:
entries 必須是 JSON 物件。entries 的 key/value 必須是非空字串。group.dicts 的順序決定。conversion_chain 步驟,不提供鎖定最終輸出。備註:OpenCC 1.3.2+ 解析器支援有限 JSONC 語法(//、/* */ 註解與尾逗號)。
若需跨實作相容,建議使用嚴格 JSON,不依賴 JSONC 擴充。
更多完整示例可見 examples/config/。該目錄僅供學習與自訂參考,不屬於官方內建
配置列表。
OpenCC 現已支援外部 C++ 分詞插件。當前第一個插件為 opencc-jieba,
可通過 s2t_jieba.json、s2tw_jieba.json、s2hk_jieba.json、
s2twp_jieba.json、tw2sp_jieba.json 等插件配置啓用。
OpenCC now supports external C++ segmentation plugins. The first plugin is
opencc-jieba, which can be enabled through plugin-backed configs such as
s2t_jieba.json, s2tw_jieba.json, s2hk_jieba.json,
s2twp_jieba.json, and tw2sp_jieba.json.
注意:
jieba 插件是可選組件,Python 套件和 Node.js 套件都不要求它。自 1.4.1 起,macOS 上的頂層 CMake 構建(含 Homebrew)預設啓用該插件(BUILD_OPENCC_JIEBA_PLUGIN=ON);其他平台與以子項目方式引入的構建預設仍為關閉。opencc-jieba 額外依賴 cppjieba 及其配套詞典資源,這些依賴僅在構建或分發該插件時需要。Notes:
jieba plugin is optional and is not required by the Python package or
the Node.js package. Since 1.4.1, top-level CMake builds on macOS (including
Homebrew) enable the plugin by default (BUILD_OPENCC_JIEBA_PLUGIN=ON);
other platforms and subproject builds keep it disabled by default.opencc-jieba additionally depends on cppjieba and its dictionary
resources. These dependencies are only needed when building or distributing
the plugin itself.g++ 4.6+ or clang 3.2+ is required.
make
build.cmd
bazel build //:opencc
make test
test.cmd
bazel test --test_output=all //src/... //data/... //python/... //test/...
make benchmark
詳情見 doc/benchmark.md 檔案。
Please update if your project is using OpenCC.
Apache License 2.0
opencc-jieba plugin.opencc-jieba 插件使用的可選依賴。自 1.4.2 起,Windows 發佈的二進位檔(可攜式 CLI 壓縮包中的執行檔,以及 npm
win32-x64 套件中的 opencc.node 與 opencc-jieba.dll)皆經過 Authenticode
簽章。
Since 1.4.2, the Windows binaries — the executables in the portable CLI zip
and the opencc.node / opencc-jieba.dll shipped in the win32-x64 npm
packages — are Authenticode-signed.
This program uses free code signing provided by SignPath.io, and a free code signing certificate by the SignPath Foundation.
本項目使用 SignPath.io 免費提供的程式碼簽章服務, 憑證由 SignPath Foundation 免費提供。
master 的提交會被簽章,簽章請求
由 維護者在發佈流程中透過 GitHub Actions 發起
(見 release-winget.yml 與
release-npm-binaries.yml)。opencc, opencc-js 与 opencc-wasm 三个 NPM packages 區別的說明
https://github.com/nk2028/opencc-js/blob/HEAD/README-zh-TW.md#%E8%88%87-opencc-npm-package-%E7%9A%84%E5%8D%80%E5%88%A5Please feel free to update this list if you have contributed OpenCC.
(top 30 of 121)
C++
70.3%
Python
9.9%
JavaScript
6.3%
Starlark
4.5%
CMake
4.4%
Shell
2.0%
Library for conversion between Traditional and Simplified Chinese
9,976
stars
1,794
commits
C++
primary language
Sep 9, 2026
updated

Open Chinese Convert (OpenCC, 開放中文轉換) is an open source project for high-quality conversion between Traditional Chinese, Simplified Chinese, Japanese Shinjitai, and regional wording across Mainland China, Taiwan, and Hong Kong. It provides dictionaries, a reusable library, conversion tools, and dictionary generation tools.
Open Chinese Convert(OpenCC,開放中文轉換) 是一個開源中文轉換項目,支持繁體中文、簡體中文、日文新字體,以及中國大陸、臺灣、香港等地區習慣用詞之間的高品質轉換,並提供詞典、可重用庫、轉換工具及詞典生成工具。
Discussion (Telegram) 討論區(Telegram): https://t.me/open_chinese_convert
詳情參閱 OpenCC 設計思想及地區詞收錄標準。 For details, see Design Principles and Regional Phrase Criteria.
brew install openccwinget install openccnpm install -g opencc 命令可安裝 OpenCC Node.js CLInpm install -g opencc opencc-jieba 命令可同時安裝 OpenCC Node.js CLI 及 Jieba 分詞插件pip install opencc 命令可安裝 Python API 及 Python CLIOpenCC 1.4.2 大幅加速了轉換熱路徑(純文字語料的整體轉換時間最多降至原本的
1/7),修復了大端序平台載入 legacy .ocd 字典得到空字典的問題,並包含一批詞庫
修正;C++ ABI 與 1.4.1 相同(SOVERSION 1.4),自 1.4.0/1.4.1 升級的下游 C++
程式無需重新連結。
opencc, opencc-jieba, and libopencc* deb packages for one architecture, with a SHA256SUMS file.https://opencc.js.org/converter?config=s2t
npm install opencc
The npm package supports Node.js >=20.17. The native addon is installed
through prebuilt @opencc/opencc-<platform>-<arch> packages covering macOS
(x64/arm64), Linux (x64/arm64), and Windows (x64). On platforms without a
prebuilt binary package, npm install builds from source with Bazel —
compiling the addon and regenerating the dictionaries — which requires a
C++ toolchain and network access; Bazel is found on PATH (bazel or
bazelisk) or fetched automatically via npx, and Bazel downloads its own
hermetic Python toolchain for the dictionary generation.
Bun and Deno can also use the npm package through their npm compatibility support.
To install the npm CLI:
npm install -g opencc
opencc -c s2t.json -i input.txt -o output.txt
The npm CLI supports basic text conversion. Plugins, --inspect,
--segmentation, and --ambiguities require the native OpenCC CLI.
import { OpenCC } from 'opencc';
async function main() {
const converter: OpenCC = new OpenCC('s2t.json');
const result: string = await converter.convertPromise('汉字');
console.log(result); // 漢字
}
Inline configurations can be passed with OpenCC.fromConfig:
import { OpenCC } from 'opencc';
const converter = OpenCC.fromConfig({
name: 'Demo Inline Config',
conversion_chain: [{
dict: { type: 'inline', entries: { '鼠标': '滑鼠' } },
}],
});
console.log(converter.convertSync('鼠标')); // 滑鼠
pip install opencc (Windows, Linux, macOS)
import opencc
converter = opencc.OpenCC('s2t.json')
converter.convert('汉字') # 漢字
The Python package also installs a basic CLI:
pip install opencc
opencc -c s2t.json -i input.txt -o output.txt
The Python CLI supports basic text conversion, --include-tofu-risk-dictionaries,
and --resource-zip. Diagnostic modes such as --inspect, --segmentation,
and --ambiguities still require the native OpenCC CLI.
#include "opencc.h"
int main() {
const opencc::SimpleConverter converter("s2t.json");
converter.Convert("汉字"); // 漢字
return 0;
}
When OpenCC is embedded in a server binary or self-contained application, the JSON config can stay small while dictionary resources are loaded from explicit resource directories:
#include <memory>
#include <vector>
#include "SimpleConverter.hpp"
int main() {
auto resources = std::make_shared<opencc::FilesystemResourceProvider>(
std::vector<std::string>{
"/opt/my-app/opencc",
"/opt/my-app/plugins/opencc-jieba",
"/usr/share/opencc",
});
const opencc::SimpleConverter converter("s2t.json", resources);
converter.Convert("汉字");
return 0;
}
FilesystemResourceProvider searches directories in order. Existing
SimpleConverter("s2t.json") and CLI behavior continue to use the config file
location, current directory, explicit paths, and installed OpenCC data directory
as before.
#include "opencc.h"
int main() {
opencc_t opencc = opencc_open("s2t.json");
const char* input = "汉字";
char* converted = opencc_convert_utf8(opencc, input, strlen(input)); // 漢字
opencc_convert_utf8_free(converted);
opencc_close(opencc);
return 0;
}
Unless otherwise noted, this section describes the native OpenCC CLI built from
the C++ toolchain. The Python and npm CLIs support basic file/stdin conversion
only, plus --include-tofu-risk-dictionaries; the Python CLI also supports
--resource-zip.
opencc --helpopencc_dict --helpOpenCC CLI supports two diagnostic modes that output JSON instead of converted text:
--segmentation — Output segmentation result only (no conversion):
echo "他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题" | opencc -c s2twp.json --segmentation
# {"input":"他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题","segments":["他","只看","了几行","日志",",就","一叶知秋",",猜到","整个","系统","是","数据库","连接池","出了","问题"]}
--inspect — Output full inspection result (segmentation + per-stage conversion + final output):
echo "他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题" | opencc -c s2twp.json --inspect
# {"input":"他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题","segments":["他","只看","了几行","日志",",就","一叶知秋",",猜到","整个","系统","是","数据库","连接池","出了","问题"],"stages":[{"index":1,"segments":["他","只看","了幾行","日誌",",就","一葉知秋",",猜到","整個","系統","是","數據庫","連接池","出了","問題"]},{"index":2,"segments":["他","只看","了幾行","日誌",",就","一葉知秋",",猜到","整個","系統","是","資料庫","連線池","出了","問題"]},{"index":3,"segments":["他","只看","了幾行","日誌",",就","一葉知秋",",猜到","整個","系統","是","資料庫","連線池","出了","問題"]}],"output":"他只看了幾行日誌,就一葉知秋,猜到整個系統是資料庫連線池出了問題"}
# Pretty-print with jq:
echo "他只看了几行日志,就一叶知秋,猜到整个系统是数据库连接池出了问题" | opencc -c s2twp.json --inspect | jq .
--ambiguities — Convert while marking every output span whose dictionary
match is one-to-many (e.g. 文丑 may be either 文丑 or 文醜), as a stream
of JSONL records:
printf '大战文丑的时候,他的头发很干燥' | opencc -c s2t.json --ambiguities
# {"def":"文丑"}
# {"lit":"大戰"}
# {"amb":{"t":"文丑","s":0}}
# {"lit":"的時候,他的頭髮很乾燥"}
# {"end":{"output_bytes":45,"ambiguities":1,"sources":1}}
Record kinds: {"def"} defines the next source index (each distinct input
word is defined once, before its first reference); {"lit"} is a literal
run of unambiguous output; {"amb":{"t","s"}} is an ambiguous span whose
output text t (the default candidate) came from the source with index s;
the final {"end"} record carries stream totals. Concatenating every lit
and amb.t in order reproduces the plain conversion output exactly.
Streaming is bounded-memory regardless of input size or line length, and
positions are implicit in record order, so consumers can rebuild offsets in
their own string-index units. Single-value conversions (e.g. 头发 →
頭髮) are not flagged.
The record stream is a machine-readable contract: every emitted line is
valid JSON, and the final {"end"} record only appears when the whole
input was processed. Input containing invalid UTF-8 aborts the stream with
an error (unlike plain conversion, which is byte-transparent), and a
missing {"end"} record means the stream is incomplete.
These modes are useful for diagnosing conversion issues:
--segmentation to verify that the input is segmented as expected.--inspect to see which conversion stage produces an unexpected result.--ambiguities to locate one-to-many conversions and resolve each
deduplicated def source to its candidates.Rules:
--segmentation, --inspect, and --ambiguities are mutually exclusive.The following ports are maintained within the OpenCC ecosystem and are generally up to date with current configuration and dictionary data.
These ports are community-maintained and may not always track upstream updates.
s2t.json Simplified Chinese to Traditional Chinese (OpenCC Standard) / 簡體 到 OpenCC 標準繁體t2s.json Traditional Chinese (OpenCC Standard) to Simplified Chinese / OpenCC 標準繁體 到 簡體s2tw.json Simplified Chinese to Traditional Chinese (Taiwan Standard) / 簡體 到 台灣正體tw2s.json Traditional Chinese (Taiwan Standard) to Simplified Chinese / 台灣正體 到 簡體s2hk.json Simplified Chinese to Traditional Chinese (Hong Kong variant) / 簡體 到 香港繁體hk2s.json Traditional Chinese (Hong Kong variant) to Simplified Chinese / 香港繁體 到 簡體s2twp.json Simplified Chinese to Traditional Chinese (Taiwan Standard, with Taiwan Phrases) / 簡體 到 台灣正體(含台灣常用詞彙)tw2sp.json Traditional Chinese (Taiwan Standard) to Simplified Chinese (Mainland China Phrases) / 台灣正體 到 簡體(含中國大陸常用詞彙)t2tw.json Traditional Chinese (OpenCC Standard) to Traditional Chinese (Taiwan Standard) / OpenCC 標準繁體 到 台灣正體tw2t.json Traditional Chinese (Taiwan Standard) to Traditional Chinese (OpenCC Standard) / 台灣正體 到 OpenCC 標準繁體t2hk.json Traditional Chinese (OpenCC Standard) to Traditional Chinese (Hong Kong variant) / OpenCC 標準繁體 到 香港繁體hk2t.json Traditional Chinese (Hong Kong variant) to Traditional Chinese (OpenCC Standard) / 香港繁體 到 OpenCC 標準繁體下列配置文件仍在開發中,歡迎貢獻新詞組:
s2hkp.json Simplified Chinese to Traditional Chinese (Hong Kong variant, with Hong Kong Phrases) / 簡體 到 香港繁體(香港常用詞彙)hk2sp.json Traditional Chinese (Hong Kong variant) to Simplified Chinese (Mainland China Phrases) / 香港繁體 到 簡體(含中國大陸常用詞彙)下列配置文件僅供探索性研究,不建議用於生產環境:
t2jp.json Old Japanese Kanji (Kyūjitai) to New Japanese Kanji (Shinjitai) / 日文舊字體 到 日文新字體jp2t.json New Japanese Kanji (Shinjitai) to Old Japanese Kanji (Kyūjitai) / 日文新字體 到 日文舊字體,並將少量日文詞組轉換爲對應中文通过环境变量OPENCC_DATA_DIR加载指定路径下的配置文件
OPENCC_DATA_DIR=/path/to/your/config/dir opencc --help
配置檔中的字典可使用 type: "inline",直接在 JSON 裡定義小型自訂詞彙,
不必修改外部字典檔。例如在 group.dicts 最前面加入覆寫規則:
{
"conversion_chain": [
{
"dict": {
"type": "group",
"dicts": [
{
"type": "inline",
"entries": {
"麦旋风": "冰炫風",
"服务器": "伺服器"
}
},
{
"type": "ocd2",
"file": "STPhrases.ocd2"
},
{
"type": "ocd2",
"file": "STCharacters.ocd2"
}
]
}
}
]
}
規則與限制:
entries 必須是 JSON 物件。entries 的 key/value 必須是非空字串。group.dicts 的順序決定。conversion_chain 步驟,不提供鎖定最終輸出。備註:OpenCC 1.3.2+ 解析器支援有限 JSONC 語法(//、/* */ 註解與尾逗號)。
若需跨實作相容,建議使用嚴格 JSON,不依賴 JSONC 擴充。
更多完整示例可見 examples/config/。該目錄僅供學習與自訂參考,不屬於官方內建
配置列表。
OpenCC 現已支援外部 C++ 分詞插件。當前第一個插件為 opencc-jieba,
可通過 s2t_jieba.json、s2tw_jieba.json、s2hk_jieba.json、
s2twp_jieba.json、tw2sp_jieba.json 等插件配置啓用。
OpenCC now supports external C++ segmentation plugins. The first plugin is
opencc-jieba, which can be enabled through plugin-backed configs such as
s2t_jieba.json, s2tw_jieba.json, s2hk_jieba.json,
s2twp_jieba.json, and tw2sp_jieba.json.
注意:
jieba 插件是可選組件,Python 套件和 Node.js 套件都不要求它。自 1.4.1 起,macOS 上的頂層 CMake 構建(含 Homebrew)預設啓用該插件(BUILD_OPENCC_JIEBA_PLUGIN=ON);其他平台與以子項目方式引入的構建預設仍為關閉。opencc-jieba 額外依賴 cppjieba 及其配套詞典資源,這些依賴僅在構建或分發該插件時需要。Notes:
jieba plugin is optional and is not required by the Python package or
the Node.js package. Since 1.4.1, top-level CMake builds on macOS (including
Homebrew) enable the plugin by default (BUILD_OPENCC_JIEBA_PLUGIN=ON);
other platforms and subproject builds keep it disabled by default.opencc-jieba additionally depends on cppjieba and its dictionary
resources. These dependencies are only needed when building or distributing
the plugin itself.g++ 4.6+ or clang 3.2+ is required.
make
build.cmd
bazel build //:opencc
make test
test.cmd
bazel test --test_output=all //src/... //data/... //python/... //test/...
make benchmark
詳情見 doc/benchmark.md 檔案。
Please update if your project is using OpenCC.
Apache License 2.0
opencc-jieba plugin.opencc-jieba 插件使用的可選依賴。自 1.4.2 起,Windows 發佈的二進位檔(可攜式 CLI 壓縮包中的執行檔,以及 npm
win32-x64 套件中的 opencc.node 與 opencc-jieba.dll)皆經過 Authenticode
簽章。
Since 1.4.2, the Windows binaries — the executables in the portable CLI zip
and the opencc.node / opencc-jieba.dll shipped in the win32-x64 npm
packages — are Authenticode-signed.
This program uses free code signing provided by SignPath.io, and a free code signing certificate by the SignPath Foundation.
本項目使用 SignPath.io 免費提供的程式碼簽章服務, 憑證由 SignPath Foundation 免費提供。
master 的提交會被簽章,簽章請求
由 維護者在發佈流程中透過 GitHub Actions 發起
(見 release-winget.yml 與
release-npm-binaries.yml)。opencc, opencc-js 与 opencc-wasm 三个 NPM packages 區別的說明
https://github.com/nk2028/opencc-js/blob/HEAD/README-zh-TW.md#%E8%88%87-opencc-npm-package-%E7%9A%84%E5%8D%80%E5%88%A5Please feel free to update this list if you have contributed OpenCC.
(top 30 of 121)
C++
70.3%
Python
9.9%
JavaScript
6.3%
Starlark
4.5%
CMake
4.4%
Shell
2.0%