kh-kim/arxiv-translator

HTML

105

318 commits

updated May 8, 2024

See the code

README

Arxiv Translation Project

이 레포는 쏟아지는 페이퍼들에 대응하기 위하여, 빠르게 Arxiv 페이퍼를 살펴볼 수 있도록 한글화된 웹페이지를 제공하는 것을 목표로 합니다. 각기 다른 형태의 PDF 파일을 번역하기 위해서, 텍스트를 추출할 때 nougat OCR 라이브러리를 활용합니다. 따라서 추출이 원활하지 않을 수 있습니다. 처음에는 Ar5iv를 번역할까 생각했지만, Ar5iv도 한달이 지나서야 페이퍼가 업데이트 되며, 최초 버전만 HTML화 하고 최종 버전은 반영되어 있지 않기 때문에, 자체적으로 내용을 추출하기로 결정하였습니다. 정확한 내용을 파악하기 위해서는 원본 페이퍼를 읽는 것을 추천합니다.

Paper List

새 창 열기가 지원되지 않습니다. 직접 새 창으로 열기를 통해 열기를 권장합니다.

ArXiv IDTitleArXivGo to
2404.19705v2When to Retrieve Teaching LLMs to Utilize Information Retrieval EffectivelyarXivpage
2404.19543RAG and RAU A Survey on Retrieval-Augmented Language Model in Natural Language ProcessingarXivpage
2404.14219v1Phi-3 Technical Report A Highly Capable Language Model Locally on Your PhonearXivpage
2404.12241v1Introducing v05 of the AI Safety Benchmark from MLCommonsarXivpage
2404.11584v1The Landscape of Emerging AI Agent Architectures for Reasoning Planning and Tool Calling A SurveyarXivpage
2404.10981v1A Survey on Retrieval-Augmented Text Generation for Large Language ModelsarXivpage
2404.10198v1How faithful are RAG models? Quantifying the tug-of-war between RAG and LLMs internal priorarXivpage
2404.10102v1Chinchilla Scaling A replication attemptarXivpage
2404.09516v1State Space Model for New-Generation Network Alternative to Transformers A SurveyarXivpage
2404.07965v1Rho-1 Not All Tokens Are What You NeedarXivpage
2404.07647v1Why do small language models underperform? Studying Language Model Saturation via the Softmax BottleneckarXivpage
2404.07503v1Best Practices and Lessons Learned on Synthetic Data for Language ModelsarXivpage
2404.07143v1Leave No Context Behind Efficient Infinite Context Transformers with Infini-attentionarXivpage
2404.06395v1MiniCPM Unveiling the Potential of Small Language Models with Scalable Training StrategiesarXivpage
2404.05875v1CodecLM Aligning Language Models with Tailored Synthetic DataarXivpage
2404.05405Physics of Language Models Part 33 Knowledge Capacity Scaling LawsarXivpage
2404.04167v3Chinese Tiny LLM Pretraining a Chinese-Centric Large Language ModelarXivpage
2404.03414v1Can Small Language Models Help Large Language Models Reason Better? LM-Guided Chain-of-ThoughtarXivpage
2404.01261v1FABLES Evaluating faithfulness and content selection in book-length summarizationarXivpage
2404.01204The Fine Line Navigating Large Language Model Pretraining with Down-streaming Capability AnalysisarXivpage
2403.19270v1sDPO Dont Use Your Data All at OncearXivpage
2403.18058v1COIG-CQIA Quality is All You Need for Chinese Instruction Fine-tuningarXivpage
2403.16971v2AIOS LLM Agent Operating SystemarXivpage
2403.16952v1Data Mixing Laws Optimizing Data Mixtures by Predicting Language Modeling PerformancearXivpage
2403.15796v2Understanding Emergent Abilities of Language Models from the Loss PerspectivearXivpage
2403.13799v1Reverse Training to Nurse the Reversal CursearXivpage
2403.13187v1Evolutionary Optimization of Model Merging RecipesarXivpage
2403.10131v1RAFT Adapting Language Model to Domain Specific RAGarXivpage
2403.09629Quiet-STaR Language Models Can Teach Themselves to Think Before SpeakingarXivpage
2403.08763Simple and Scalable Strategies to Continually Pre-train Large Language ModelsarXivpage
2403.06634Stealing Part of a Production Language ModelarXivpage
2403.06563v1Unraveling the Mystery of Scaling Laws Part IarXivpage
2403.04706v1Common 7B Language Models Already Possess Strong Math CapabilitiesarXivpage
2403.04652v1Yi Open Foundation Models by 01AIarXivpage
2403.03883v2SaulLM-7B A pioneering Large Language Model for LawarXivpage
2403.02178v1Masked Thought Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language ModelsarXivpage
2403.01432v2Fine Tuning vs Retrieval Augmented Generation for Less Popular KnowledgearXivpage
2402.18815v1How do Large Language Models Handle Multilingualism?arXivpage
2402.18563v1Approaching Human-Level Forecasting with Language ModelsarXivpage
2402.16837v1Do Large Language Models Latently Perform Multi-Hop Reasoning?arXivpage
2402.16819v2Nemotron-4 15B Technical ReportarXivpage
2402.14714v1Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language ModelsarXivpage
2402.12847v1Instruction-tuned Language Models are Better Knowledge LearnersarXivpage
2402.08939v1Premise Order Matters in Reasoning with Large Language ModelsarXivpage
2402.07043v1A Tale of Tails Model Collapse as a Change of Scaling LawsarXivpage
2402.06196v2Large Language Models A SurveyarXivpage
2402.05120v1More Agents Is All You NeedarXivpage
2402.00838v3OLMo Accelerating the Science of Language ModelsarXivpage
2401.16380v1Rephrasing the Web A Recipe for Compute and Data-Efficient Language ModelingarXivpage
2401.10225v1ChatQA Building GPT-4 Level Conversational QA ModelsarXivpage
2401.08417v3Contrastive Preference Optimization Pushing the Boundaries of LLM Performance in Machine TranslationarXivpage
2401.05654v1Towards Conversational Diagnostic AIarXivpage
2401.03129v1Examining Forgetting in Continual Pre-training of Aligned Large Language ModelsarXivpage
2401.01055v2LLaMA Beyond English An Empirical Study on Language Capability TransferarXivpage
2312.05934v3Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsarXivpage
2311.13647Language Model InversionarXivpage
2311.08545Efficient Continual Pre-training for Building Domain Specific Large Language ModelsarXivpage
2310.11511Self-RAG Learning to Retrieve Generate and Critique through Self-ReflectionarXivpage
2310.08754v4Tokenizer Choice For LLM Training Negligible or Crucial?arXivpage
2310.04799v2Chat Vector A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New LanguagesarXivpage
2309.15402A Survey of Chain of Thought Reasoning Advances Frontiers and FuturearXivpage
2309.12288The Reversal Curse LLMs trained on A is B fail to learn B is AarXivpage
2308.12284D4 Improving LLM Pretraining via Document De-Duplication and DiversificationarXivpage
2308.11432v5A Survey on Large Language Model based Autonomous AgentsarXivpage
2308.09583WizardMath Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-InstructarXivpage
2306.08568WizardCoder Empowering Code Large Language Models with Evol-InstructarXivpage
2306.01116The RefinedWeb Dataset for Falcon LLM Outperforming Curated Corpora with Web Data and Web Data OnlyarXivpage
2305.18290v2Direct Preference Optimization Your Language Model is Secretly a Reward ModelarXivpage
2304.12244WizardLM Empowering Large Language Models to Follow Complex InstructionsarXivpage
2304.08177v3Efficient and Effective Text Encoding for Chinese LLaMA and AlpacaarXivpage
2303.18223A Survey of Large Language ModelsarXivpage
2212.10560Self-Instruct Aligning Language Models with Self-Generated InstructionsarXivpage
2110.03215Towards Continual Knowledge Learning of Language ModelsarXivpage
2107.06499Deduplicating Training Data Makes Language Models BetterarXivpage

Procedure

Arxiv 페이퍼를 번역하기 위해서 총 4단계를 거칩니다.

ArXiv Paper Download

Arxiv는 wget 등의 명령어를 통해서 pdf 파일을 다운로드 받을 수 없게 하였습니다. 아마도 무분별한 scrapping에 대응하기 위한 것으로 생각됩니다. 따라서 pdf 파일을 다운로드 받기 위해서 arxiv-dl 패키지를 활용합니다.

PDF to Markdown

Nougat OCR을 활용하여 Mathpix Markdown 파일로 변환합니다.

Translation

자체 번역 모델을 활용하여 번역을 수행합니다. 다음과 같이 페이퍼의 번역을 위해 사용된 번역기의 성능(초록색)은 DeepL과 Google, Naver의 중간쯤에 위치합니다.

NMT Evaluation Results

Markdown to HTML

Mathpix Markdown을 HTML로 변환합니다. 변환 방법은 여기에 설명되어 있습니다. 그리고 저장된 github에 push되어 저장된 HTML 파일을 githack.com을 통해 렌더링하도록 합니다.

Future Work

페이퍼 중간의 이미지들은 Nougat OCR에서 추출해주지 않기 때문에 빠져 있습니다. 따라서 이미지도 함께 포함하여 결과물을 만들어내도록 하고자 합니다.

Contact

Kim Ki Hyun pointzz.ki@gmail.com

Contributors

kh-kim

318 commits

kh-kim/arxiv-translator

HTML

105

318 commits

updated May 8, 2024

See the code

README

Arxiv Translation Project

이 레포는 쏟아지는 페이퍼들에 대응하기 위하여, 빠르게 Arxiv 페이퍼를 살펴볼 수 있도록 한글화된 웹페이지를 제공하는 것을 목표로 합니다. 각기 다른 형태의 PDF 파일을 번역하기 위해서, 텍스트를 추출할 때 nougat OCR 라이브러리를 활용합니다. 따라서 추출이 원활하지 않을 수 있습니다. 처음에는 Ar5iv를 번역할까 생각했지만, Ar5iv도 한달이 지나서야 페이퍼가 업데이트 되며, 최초 버전만 HTML화 하고 최종 버전은 반영되어 있지 않기 때문에, 자체적으로 내용을 추출하기로 결정하였습니다. 정확한 내용을 파악하기 위해서는 원본 페이퍼를 읽는 것을 추천합니다.

Paper List

새 창 열기가 지원되지 않습니다. 직접 새 창으로 열기를 통해 열기를 권장합니다.

ArXiv IDTitleArXivGo to
2404.19705v2When to Retrieve Teaching LLMs to Utilize Information Retrieval EffectivelyarXivpage
2404.19543RAG and RAU A Survey on Retrieval-Augmented Language Model in Natural Language ProcessingarXivpage
2404.14219v1Phi-3 Technical Report A Highly Capable Language Model Locally on Your PhonearXivpage
2404.12241v1Introducing v05 of the AI Safety Benchmark from MLCommonsarXivpage
2404.11584v1The Landscape of Emerging AI Agent Architectures for Reasoning Planning and Tool Calling A SurveyarXivpage
2404.10981v1A Survey on Retrieval-Augmented Text Generation for Large Language ModelsarXivpage
2404.10198v1How faithful are RAG models? Quantifying the tug-of-war between RAG and LLMs internal priorarXivpage
2404.10102v1Chinchilla Scaling A replication attemptarXivpage
2404.09516v1State Space Model for New-Generation Network Alternative to Transformers A SurveyarXivpage
2404.07965v1Rho-1 Not All Tokens Are What You NeedarXivpage
2404.07647v1Why do small language models underperform? Studying Language Model Saturation via the Softmax BottleneckarXivpage
2404.07503v1Best Practices and Lessons Learned on Synthetic Data for Language ModelsarXivpage
2404.07143v1Leave No Context Behind Efficient Infinite Context Transformers with Infini-attentionarXivpage
2404.06395v1MiniCPM Unveiling the Potential of Small Language Models with Scalable Training StrategiesarXivpage
2404.05875v1CodecLM Aligning Language Models with Tailored Synthetic DataarXivpage
2404.05405Physics of Language Models Part 33 Knowledge Capacity Scaling LawsarXivpage
2404.04167v3Chinese Tiny LLM Pretraining a Chinese-Centric Large Language ModelarXivpage
2404.03414v1Can Small Language Models Help Large Language Models Reason Better? LM-Guided Chain-of-ThoughtarXivpage
2404.01261v1FABLES Evaluating faithfulness and content selection in book-length summarizationarXivpage
2404.01204The Fine Line Navigating Large Language Model Pretraining with Down-streaming Capability AnalysisarXivpage
2403.19270v1sDPO Dont Use Your Data All at OncearXivpage
2403.18058v1COIG-CQIA Quality is All You Need for Chinese Instruction Fine-tuningarXivpage
2403.16971v2AIOS LLM Agent Operating SystemarXivpage
2403.16952v1Data Mixing Laws Optimizing Data Mixtures by Predicting Language Modeling PerformancearXivpage
2403.15796v2Understanding Emergent Abilities of Language Models from the Loss PerspectivearXivpage
2403.13799v1Reverse Training to Nurse the Reversal CursearXivpage
2403.13187v1Evolutionary Optimization of Model Merging RecipesarXivpage
2403.10131v1RAFT Adapting Language Model to Domain Specific RAGarXivpage
2403.09629Quiet-STaR Language Models Can Teach Themselves to Think Before SpeakingarXivpage
2403.08763Simple and Scalable Strategies to Continually Pre-train Large Language ModelsarXivpage
2403.06634Stealing Part of a Production Language ModelarXivpage
2403.06563v1Unraveling the Mystery of Scaling Laws Part IarXivpage
2403.04706v1Common 7B Language Models Already Possess Strong Math CapabilitiesarXivpage
2403.04652v1Yi Open Foundation Models by 01AIarXivpage
2403.03883v2SaulLM-7B A pioneering Large Language Model for LawarXivpage
2403.02178v1Masked Thought Simply Masking Partial Reasoning Steps Can Improve Mathematical Reasoning Learning of Language ModelsarXivpage
2403.01432v2Fine Tuning vs Retrieval Augmented Generation for Less Popular KnowledgearXivpage
2402.18815v1How do Large Language Models Handle Multilingualism?arXivpage
2402.18563v1Approaching Human-Level Forecasting with Language ModelsarXivpage
2402.16837v1Do Large Language Models Latently Perform Multi-Hop Reasoning?arXivpage
2402.16819v2Nemotron-4 15B Technical ReportarXivpage
2402.14714v1Efficient and Effective Vocabulary Expansion Towards Multilingual Large Language ModelsarXivpage
2402.12847v1Instruction-tuned Language Models are Better Knowledge LearnersarXivpage
2402.08939v1Premise Order Matters in Reasoning with Large Language ModelsarXivpage
2402.07043v1A Tale of Tails Model Collapse as a Change of Scaling LawsarXivpage
2402.06196v2Large Language Models A SurveyarXivpage
2402.05120v1More Agents Is All You NeedarXivpage
2402.00838v3OLMo Accelerating the Science of Language ModelsarXivpage
2401.16380v1Rephrasing the Web A Recipe for Compute and Data-Efficient Language ModelingarXivpage
2401.10225v1ChatQA Building GPT-4 Level Conversational QA ModelsarXivpage
2401.08417v3Contrastive Preference Optimization Pushing the Boundaries of LLM Performance in Machine TranslationarXivpage
2401.05654v1Towards Conversational Diagnostic AIarXivpage
2401.03129v1Examining Forgetting in Continual Pre-training of Aligned Large Language ModelsarXivpage
2401.01055v2LLaMA Beyond English An Empirical Study on Language Capability TransferarXivpage
2312.05934v3Fine-Tuning or Retrieval? Comparing Knowledge Injection in LLMsarXivpage
2311.13647Language Model InversionarXivpage
2311.08545Efficient Continual Pre-training for Building Domain Specific Large Language ModelsarXivpage
2310.11511Self-RAG Learning to Retrieve Generate and Critique through Self-ReflectionarXivpage
2310.08754v4Tokenizer Choice For LLM Training Negligible or Crucial?arXivpage
2310.04799v2Chat Vector A Simple Approach to Equip LLMs with Instruction Following and Model Alignment in New LanguagesarXivpage
2309.15402A Survey of Chain of Thought Reasoning Advances Frontiers and FuturearXivpage
2309.12288The Reversal Curse LLMs trained on A is B fail to learn B is AarXivpage
2308.12284D4 Improving LLM Pretraining via Document De-Duplication and DiversificationarXivpage
2308.11432v5A Survey on Large Language Model based Autonomous AgentsarXivpage
2308.09583WizardMath Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-InstructarXivpage
2306.08568WizardCoder Empowering Code Large Language Models with Evol-InstructarXivpage
2306.01116The RefinedWeb Dataset for Falcon LLM Outperforming Curated Corpora with Web Data and Web Data OnlyarXivpage
2305.18290v2Direct Preference Optimization Your Language Model is Secretly a Reward ModelarXivpage
2304.12244WizardLM Empowering Large Language Models to Follow Complex InstructionsarXivpage
2304.08177v3Efficient and Effective Text Encoding for Chinese LLaMA and AlpacaarXivpage
2303.18223A Survey of Large Language ModelsarXivpage
2212.10560Self-Instruct Aligning Language Models with Self-Generated InstructionsarXivpage
2110.03215Towards Continual Knowledge Learning of Language ModelsarXivpage
2107.06499Deduplicating Training Data Makes Language Models BetterarXivpage

Procedure

Arxiv 페이퍼를 번역하기 위해서 총 4단계를 거칩니다.

ArXiv Paper Download

Arxiv는 wget 등의 명령어를 통해서 pdf 파일을 다운로드 받을 수 없게 하였습니다. 아마도 무분별한 scrapping에 대응하기 위한 것으로 생각됩니다. 따라서 pdf 파일을 다운로드 받기 위해서 arxiv-dl 패키지를 활용합니다.

PDF to Markdown

Nougat OCR을 활용하여 Mathpix Markdown 파일로 변환합니다.

Translation

자체 번역 모델을 활용하여 번역을 수행합니다. 다음과 같이 페이퍼의 번역을 위해 사용된 번역기의 성능(초록색)은 DeepL과 Google, Naver의 중간쯤에 위치합니다.

NMT Evaluation Results

Markdown to HTML

Mathpix Markdown을 HTML로 변환합니다. 변환 방법은 여기에 설명되어 있습니다. 그리고 저장된 github에 push되어 저장된 HTML 파일을 githack.com을 통해 렌더링하도록 합니다.

Future Work

페이퍼 중간의 이미지들은 Nougat OCR에서 추출해주지 않기 때문에 빠져 있습니다. 따라서 이미지도 함께 포함하여 결과물을 만들어내도록 하고자 합니다.

Contact

Kim Ki Hyun pointzz.ki@gmail.com

Contributors

kh-kim

318 commits

Languages

HTML

92.0%

Mermaid

8.0%