AI Frontier

AI FRONTIER / 原始报告

持续追踪前沿 AI,从原始报告走向技术洞察。

原始报告,前沿模型,技术洞察。按厂商与类别探索 AI 的演进。

探索报告

原始文档 · 清晰分类 · 来源可追溯

模型核查:2026-10-03
385原始报告
18厂商
125前沿模型
2022 — 2026原始文档 · 清晰分类 · 来源可追溯

追踪每一家厂商

选择厂商,直达原始报告。

模型核查:2026-10-03
385 条结果
Google / DeepMind
Technical Report

WeatherNext 3: Increasing Resolution and Performance of Global Weather Models with Raw Observations

2026-09-03 · arxiv-v1 · 43 页

Alibaba / Qwen / Wan
Technical Report

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

2026-08-31 · arxiv-v1 · 40 页

Alibaba / Qwen / Wan
Technical Report

Wan-Animate-2: Pushing the Application Boundaries of Character Animation

2026-08-06 · arxiv-v2 · 14 页

Tencent / Hunyuan
Technical Report

WorldClaw: Agentic 3D Open-World Generation at Scale

2026-08-05 · arxiv-v1 · 38 页

Tencent / Hunyuan
Technical Report

Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing

2026-08-03 · arxiv-v3 · 33 页

Alibaba / Qwen / Wan
Technical Report

Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

2026-07-30 · arxiv-v1 · 56 页

Alibaba / Qwen / Wan
Technical Report

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

2026-07-10 · arxiv-v3 · 17 页

NVIDIA
Technical Report

Unified Audio Intelligence Without Regressing on Text Intelligence

2026-07-06 · arxiv-v2 · 41 页

Tencent / Hunyuan
Technical Report

HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better

2026-07-06 · arxiv-v3 · 41 页

NVIDIA
Technical Report

Nemotron-Labs-3-Puzzle-75B-A9B: Compressing Hybrid MoE LLMs

2026-07-05 · arxiv-v2 · 25 页

Alibaba / Qwen / Wan
Technical Report

Qwen-RobotNav Technical Report: A Scalable Navigation Model Designed for an Agentic Navigation System

2026-06-16 · arxiv-v3 · 37 页

Alibaba / Qwen / Wan
Technical Report

Qwen-RobotManip Technical Report: Alignment Unlocks Scale for Robotic Manipulation Foundation Models

2026-06-16 · arxiv-v2 · 44 页

NVIDIA
Technical Report

Cosmos 3: Omnimodal World Models for Physical AI

2026-06-01 · 139 页

Alibaba / Qwen / Wan
Technical Report

Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments

2026-05-28 · arxiv-v2 · 34 页

MiniMax
Technical Report

The MiniMax-M2 Series: Mini Activations Unleashing Max Real-World Intelligence

2026-05-26 · arxiv-v2 · 35 页

Tencent / Hunyuan
Technical Report

Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild

2026-05-21 · arxiv-v2 · 15 页

Alibaba / Qwen / Wan
Technical Report

Qwen-Image-2.0 Technical Report

2026-05-11 · arxiv-v1 · 30 页

DeepSeek
Technical Report

DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence

2026-04-26 · arxiv-v1 · 58 页

ByteDance / Seed
Technical Report

Seed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content Generation

2026-04-22 · arxiv-v1 · 18 页

Alibaba / Qwen / Wan
Technical Report

Qwen3.5-Omni Technical Report

2026-04-17 · arxiv-v2 · 28 页

ByteDance / Seed
Technical Report

Seedance 2.0: Advancing Video Generation for World Complexity

2026-04-15 · arxiv-v1 · 26 页

NVIDIA
Technical Report

Nemotron-Cascade 2: Post-Training LLMs with Cascade RL and Multi-Domain On-Policy Distillation

2026-03-19 · arxiv-v2 · 63 页

Meta
Technical Report

V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning

2026-03-15 · arxiv-v3 · 37 页

Cohere
Technical Report

Tiny Aya: Bridging Scale and Multilingual Depth

2026-03-12 · arxiv-v1 · 50 页

智谱 / Z.ai
Technical Report

GLM-5: from Vibe Coding to Agentic Engineering

2026-02-17 · arxiv-v2 · 40 页

StepFun
Technical Report

Step 3.5 Flash: Open Frontier-Level Intelligence with 11B Active Parameters

2026-02-11 · arxiv-v2 · 67 页

Alibaba / Qwen / Wan
Technical Report

Qwen3-ASR Technical Report

2026-01-29 · arxiv-v2 · 16 页

Alibaba / Qwen / Wan
Technical Report

Qwen3-TTS Technical Report

2026-01-22 · arxiv-v1 · 14 页

Alibaba / Qwen / Wan
Technical Report

Qwen3-VL-Embedding and Qwen3-VL-Reranker Technical Report

2026-01-08 · 23 页

StepFun
Technical Report

Step-DeepResearch Technical Report

2025-12-23 · arxiv-v4 · 27 页

Meta
Technical Report

Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning

2025-12-22 · arxiv-v1 · 38 页

Meta
Technical Report

SAM Audio: Segment Anything in Audio

2025-12-19 · arxiv-v1 · 57 页

ByteDance / Seed
Model Card

Seed1.8 Model Card: Towards Generalized Real-World Agency

2025-12-17 · 48 页

Alibaba / Qwen / Wan
Technical Report

Qwen-Image-Layered: Towards Inherent Editability via Layer Decomposition

2025-12-17 · arxiv-v1 · 12 页

Google / DeepMind
Technical Report

T5Gemma 2: Seeing, Reading, and Understanding Longer

2025-12-16 · arxiv-v2 · 13 页

MiniMax
Technical Report

Towards Scalable Pre-training of Visual Tokenizers for Generation

2025-12-15 · arxiv-v2 · 16 页

ByteDance / Seed
Model Card

Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

2025-12-15 · arxiv-v3 · 11 页

NVIDIA
Technical Report

Nemotron-Cascade: Scaling Cascaded Reinforcement Learning for General-Purpose Reasoning Models

2025-12-15 · arxiv-v2 · 60 页

NVIDIA
Technical Report

NVIDIA Nemotron 3: Efficient and Open Intelligence

2025-12-15 · 13 页

Alibaba / Qwen / Wan
Technical Report

Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance

2025-12-09 · arxiv-v1 · 22 页

DeepSeek
Technical Report

DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

2025-12-02 · arxiv-v1 · 23 页

StepFun
Technical Report

ReasonEdit: Towards Reasoning-Enhanced Image Editing Models

2025-11-27 · arxiv-v2 · 18 页

DeepSeek
Technical Report

DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning

2025-11-27 · 19 页

Alibaba / Qwen / Wan
Technical Report

Qwen3-VL Technical Report

2025-11-26 · arxiv-v2 · 42 页

Tencent / Hunyuan
Technical Report

HunyuanOCR Technical Report

2025-11-24 · arxiv-v2 · 36 页

Meta
Technical Report

SAM 3: Segment Anything with Concepts

2025-11-20 · arxiv-v2 · 78 页

OpenAI
System Card

GPT-5.1 Instant and GPT-5.1 Thinking System Card Addendum

2025-11-12 · 5 页

StepFun
Technical Report

Step-Audio-EditX Technical Report

2025-11-05 · arxiv-v2 · 13 页

Cohere
Technical Report

Command-A-Translate: Raising the Bar of Machine Translation with Difficulty Filtering

2025-11 · 11 页

Moonshot AI / Kimi
Technical Report

Kimi Linear: An Expressive, Efficient Attention Architecture

2025-10-30 · arxiv-v2 · 28 页

NVIDIA
Technical Report

Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail

2025-10-30 · arxiv-v2 · 42 页

OpenAI
Technical Report

Technical Report: Performance and baseline evaluations of gpt-oss-safeguard-120b and gpt-oss-safeguard-20b

2025-10-29 · 10 页

Alibaba / Qwen / Wan
Technical Report

Tongyi DeepResearch Technical Report

2025-10-28 · arxiv-v3 · 23 页

OpenAI
System Card

Addendum to GPT-5 System Card: Sensitive Conversations

2025-10-27 · 4 页

DeepSeek
Technical Report

DeepSeek-OCR: Contexts Optical Compression

2025-10-21 · arxiv-v1 · 22 页

Alibaba / Qwen / Wan
Technical Report

Qwen3Guard Technical Report

2025-10-16 · arxiv-v1 · 28 页

Moonshot AI / Kimi
Technical Report

Kimi-Dev: Agentless Training as Skill Prior for SWE-Agents

2025-09-27 · arxiv-v3 · 68 页

ByteDance / Seed
Technical Report

Seedream 4.0: Toward Next-generation Multimodal Image Generation

2025-09-24 · arxiv-v3 · 19 页

Google / DeepMind
Technical Report

EmbeddingGemma: Powerful and Lightweight Text Representations

2025-09-24 · arxiv-v3 · 18 页

Alibaba / Qwen / Wan
Technical Report

Qwen3-Omni Technical Report

2025-09-22 · arxiv-v1 · 25 页

Alibaba / Qwen / Wan
Technical Report

Wan-Animate: Unified Character Animation and Replacement with Holistic Replication

2025-09-17 · arxiv-v1 · 17 页

Tencent / Hunyuan
Technical Report

Hunyuan-MT Technical Report

2025-09-05 · arxiv-v2 · 18 页

Alibaba / Qwen / Wan
Technical Report

Wan-S2V: Audio-Driven Cinematic Video Generation

2025-08-26 · arxiv-v1 · 11 页

NVIDIA
Technical Report

Jet-Nemotron: Efficient Language Model with Post Neural Architecture Search

2025-08-21 · arxiv-v3 · 20 页

NVIDIA
Technical Report

NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

2025-08-20 · arxiv-v4 · 43 页

StepFun
Technical Report

NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

2025-08-14 · arxiv-v2 · 25 页

智谱 / Z.ai
Technical Report

GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

2025-08-08 · arxiv-v1 · 26 页

ByteDance / Seed
Technical Report

Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference

2025-08-04 · arxiv-v1 · 11 页

Alibaba / Qwen / Wan
Technical Report

Qwen-Image Technical Report

2025-08-04 · arxiv-v1 · 46 页

Moonshot AI / Kimi
Technical Report

Kimi K2: Open Agentic Intelligence

2025-07-28 · arxiv-v2 · 32 页

StepFun
Technical Report

Step-3 is Large yet Affordable: Model-system Co-design for Cost-effective Decoding

2025-07-25 · arxiv-v1 · 18 页

Microsoft / Phi / MAI
Technical Report

Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long Generation

2025-07-09 · arxiv-v3 · 35 页

Google / DeepMind
Technical Report

MedGemma Technical Report

2025-07-07 · arxiv-v4 · 60 页

Google / DeepMind
Technical Report

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

2025-07-07 · arxiv-v6 · 73 页

Amazon / Nova
System Card

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework

2025-07-07 · arxiv-v1 · 10 页

智谱 / Z.ai
Technical Report

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

2025-07-01 · arxiv-v6 · 42 页

Tencent / Hunyuan
Technical Report

Hunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR Material

2025-06-18 · arxiv-v1 · 14 页

MiniMax
Technical Report

MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

2025-06-16 · arxiv-v1 · 22 页

NVIDIA
Technical Report

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy

2025-06-16 · arxiv-v1 · 23 页

Meta
Technical Report

V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning

2025-06-11 · arxiv-v1 · 48 页

StepFun
Technical Report

Step-Audio-AQAA: a Fully End-to-End Expressive Large Audio Language Model

2025-06-10 · arxiv-v2 · 12 页

ByteDance / Seed
Technical Report

Seedance 1.0: Exploring the Boundaries of Video Generation Models

2025-06-10 · arxiv-v2 · 26 页

Alibaba / Qwen / Wan
Technical Report

Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

2025-06-05 · arxiv-v3 · 14 页

NVIDIA
Technical Report

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

2025-05-22 · arxiv-v3 · 23 页

OpenAI
System Card

Addendum to OpenAI o3 and o4-mini system card: Codex

2025-05-16 · 8 页

Alibaba / Qwen / Wan
Technical Report

WorldPM: Scaling Human Preference Modeling

2025-05-15 · arxiv-v2 · 34 页

Alibaba / Qwen / Wan
Technical Report

Qwen3 Technical Report

2025-05-14 · arxiv-v1 · 35 页

Cohere
Technical Report

Aya Vision: Advancing the Frontier of Multilingual Multimodality

2025-05-13 · arxiv-v1 · 76 页

MiniMax
Technical Report

MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

2025-05-12 · arxiv-v1 · 20 页

ByteDance / Seed
Technical Report

Seed1.5-VL Technical Report

2025-05-11 · arxiv-v1 · 77 页

NVIDIA
Technical Report

Llama-Nemotron: Efficient Reasoning Models

2025-05-02 · arxiv-v5 · 24 页

Microsoft / Phi / MAI
Technical Report

Phi-4-reasoning Technical Report

2025-04-30 · arxiv-v1 · 33 页

Meta
Technical Report

Perception Encoder: The best visual embeddings are not at the output of the network

2025-04-17 · arxiv-v2 · 44 页

ByteDance / Seed
Technical Report

Seedream 3.0 Technical Report

2025-04-15 · arxiv-v3 · 22 页

Moonshot AI / Kimi
Technical Report

Kimina-Prover Preview: Towards Large Formal Reasoning Models with Reinforcement Learning

2025-04-15 · arxiv-v1 · 24 页

ByteDance / Seed
Technical Report

Seed1.5-Thinking: Advancing Superb Reasoning Models with Reinforcement Learning

2025-04-10 · arxiv-v3 · 19 页

Moonshot AI / Kimi
Technical Report

Kimi-VL Technical Report

2025-04-10 · arxiv-v3 · 24 页

Google / DeepMind
Technical Report

TxGemma: Efficient and Agentic LLMs for Therapeutics

2025-04-08 · arxiv-v1 · 58 页

Google / DeepMind
Technical Report

Encoder-Decoder Gemma: Improving the Quality-Efficiency Trade-Off via Adaptation

2025-04-08 · arxiv-v1 · 12 页

Amazon / Nova
Technical Report

Amazon Nova Sonic: Technical Report and Model Card

2025-04-08 · 11 页

NVIDIA
Technical Report

Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

2025-04-04 · arxiv-v4 · 35 页

Cohere
Technical Report

Command A: An Enterprise-Ready Large Language Model

2025-04-01 · arxiv-v2 · 55 页

Alibaba / Qwen / Wan
Technical Report

Wan: Open and Advanced Large-Scale Video Generative Models

2025-03-26 · arxiv-v2 · 60 页

Alibaba / Qwen / Wan
Technical Report

Qwen2.5-Omni Technical Report

2025-03-26 · arxiv-v1 · 20 页

Google / DeepMind
Technical Report

Gemma 3 Technical Report

2025-03-25 · arxiv-v1 · 25 页

Google / DeepMind
Technical Report

Gemini Robotics: Bringing AI into the Physical World

2025-03-25 · arxiv-v1 · 64 页

OpenAI
System Card

Addendum to GPT-4o System Card: Native image generation

2025-03-25 · 13 页

NVIDIA
Technical Report

GR00T N1: An Open Foundation Model for Generalist Humanoid Robots

2025-03-18 · arxiv-v2 · 36 页

NVIDIA
Technical Report

Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

2025-03-18 · arxiv-v3 · 36 页

StepFun
Technical Report

Step-Video-TI2V Technical Report: A State-of-the-Art Text-Driven Image-to-Video Generation Model

2025-03-14 · arxiv-v1 · 7 页

Alibaba / Qwen / Wan
Technical Report

VACE: All-in-One Video Creation and Editing

2025-03-10 · arxiv-v2 · 17 页

ByteDance / Seed
Technical Report

Seedream 2.0: A Native Chinese-English Bilingual Image Generation Foundation Model

2025-03-10 · arxiv-v1 · 33 页

Microsoft / Phi / MAI
Technical Report

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

2025-03-03 · arxiv-v2 · 39 页

Moonshot AI / Kimi
Technical Report

Muon is Scalable for LLM Training

2025-02-24 · arxiv-v1 · 19 页

Alibaba / Qwen / Wan
Technical Report

Qwen2.5-VL Technical Report

2025-02-19 · arxiv-v1 · 23 页

StepFun
Technical Report

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

2025-02-17 · arxiv-v2 · 25 页

DeepSeek
Technical Report

Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

2025-01-29 · arxiv-v1 · 13 页

Moonshot AI / Kimi
Technical Report

Kimi k1.5: Scaling Reinforcement Learning with LLMs

2025-01-22 · arxiv-v4 · 25 页

DeepSeek
Technical Report

DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

2025-01-22 · arxiv-v2 · 86 页

Tencent / Hunyuan
Technical Report

Hunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets Generation

2025-01-21 · arxiv-v5 · 28 页

MiniMax
Technical Report

MiniMax-01: Scaling Foundation Models with Lightning Attention

2025-01-14 · arxiv-v1 · 68 页

NVIDIA
Technical Report

Cosmos World Foundation Model Platform for Physical AI

2025-01-07 · arxiv-v3 · 75 页

Alibaba / Qwen / Wan
Technical Report

Qwen2.5 Technical Report

2024-12-19 · arxiv-v2 · 26 页

DeepSeek
Technical Report

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

2024-12-13 · arxiv-v1 · 28 页

Microsoft / Phi / MAI
Technical Report

Phi-4 Technical Report

2024-12-12 · arxiv-v1 · 36 页

Meta
Technical Report

Large Concept Models: Language Modeling in a Sentence Representation Space

2024-12-11 · arxiv-v2 · 49 页

Cohere
Technical Report

Aya Expanse: Combining Research Breakthroughs for a New Multilingual Frontier

2024-12-05 · arxiv-v1 · 19 页

Google / DeepMind
Technical Report

PaliGemma 2: A Family of Versatile VLMs for Transfer

2024-12-04 · arxiv-v1 · 31 页

Amazon / Nova
Technical Report

The Amazon Nova Family of Models: Technical Report and Model Card

2024-12-03 · 48 页

Tencent / Hunyuan
Technical Report

HunyuanVideo: A Systematic Framework For Large Video Generative Models

2024-12-03 · arxiv-v6 · 35 页

智谱 / Z.ai
Technical Report

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

2024-12-03 · arxiv-v1 · 14 页

Meta
Technical Report

Llama Guard 3-1B-INT4: Compact and Efficient Safeguard for Human-AI Conversations

2024-11-18 · arxiv-v1 · 9 页

DeepSeek
Technical Report

JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

2024-11-12 · arxiv-v2 · 25 页

Tencent / Hunyuan
Technical Report

Hunyuan-Large: An Open-Source MoE Model with 52 Billion Activated Parameters by Tencent

2024-11-04 · arxiv-v3 · 18 页

Meta
Technical Report

Movie Gen: A Cast of Media Foundation Models

2024-10-17 · arxiv-v2 · 96 页

DeepSeek
Technical Report

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

2024-10-17 · arxiv-v1 · 24 页

Anthropic
Model Card

Model Card Addendum: Claude 3.5 Haiku and Upgraded Claude 3.5 Sonnet

2024-10 · 14 页

Alibaba / Qwen / Wan
Technical Report

Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

2024-09-18 · arxiv-v1 · 39 页

Alibaba / Qwen / Wan
Technical Report

Qwen2.5-Coder Technical Report

2024-09-18 · arxiv-v3 · 32 页

Alibaba / Qwen / Wan
Technical Report

Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

2024-09-18 · arxiv-v2 · 52 页

智谱 / Z.ai
Technical Report

CogVLM2: Visual Language Models for Image and Video Understanding

2024-08-29 · arxiv-v1 · 27 页

DeepSeek
Technical Report

DeepSeek-Prover-V1.5: Harnessing Proof Assistant Feedback for Reinforcement Learning and Monte-Carlo Tree Search

2024-08-15 · arxiv-v1 · 28 页

智谱 / Z.ai
Technical Report

CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

2024-08-12 · arxiv-v3 · 30 页

Meta
Technical Report

SAM 2: Segment Anything in Images and Videos

2024-08-01 · arxiv-v2 · 42 页

Alibaba / Qwen / Wan
Technical Report

Qwen2-Audio Technical Report

2024-07-15 · arxiv-v1 · 16 页

Alibaba / Qwen / Wan
Technical Report

Qwen2 Technical Report

2024-07-15 · arxiv-v4 · 26 页

Google / DeepMind
Technical Report

PaliGemma: A versatile 3B VLM for transfer

2024-07-10 · arxiv-v2 · 59 页

智谱 / Z.ai
Technical Report

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

2024-06-18 · arxiv-v2 · 19 页

DeepSeek
Technical Report

DeepSeek-Coder-V2: Breaking the Barrier of Closed-Source Models in Code Intelligence

2024-06-17 · arxiv-v1 · 19 页

Google / DeepMind
Technical Report

CodeGemma: Open Code Models Based on Gemma

2024-06-17 · arxiv-v2 · 11 页

DeepSeek
Technical Report

DeepSeek-Prover: Advancing Theorem Proving in LLMs through Large-Scale Synthetic Data

2024-05-23 · arxiv-v1 · 17 页

Cohere
Technical Report

Aya 23: Open Weight Releases to Further Multilingual Progress

2024-05-23 · arxiv-v2 · 27 页

Meta
Technical Report

Chameleon: Mixed-Modal Early-Fusion Foundation Models

2024-05-16 · arxiv-v2 · 27 页

Tencent / Hunyuan
Technical Report

Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

2024-05-14 · arxiv-v1 · 25 页

智谱 / Z.ai
Technical Report

Inf-DiT: Upsampling Any-Resolution Image with Memory-Efficient Diffusion Transformer

2024-05-07 · arxiv-v2 · 24 页

DeepSeek
Technical Report

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

2024-05-07 · arxiv-v5 · 52 页

Microsoft / Phi / MAI
Technical Report

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

2024-04-22 · arxiv-v4 · 24 页

Google / DeepMind
Technical Report

Gemma: Open Models Based on Gemini Research and Technology

2024-03-13 · arxiv-v4 · 17 页

DeepSeek
Technical Report

DeepSeek-VL: Towards Real-World Vision-Language Understanding

2024-03-08 · arxiv-v2 · 33 页

智谱 / Z.ai
Technical Report

CogView3: Finer and Faster Text-to-Image Generation via Relay Diffusion

2024-03-08 · arxiv-v1 · 21 页

Cohere
Technical Report

Aya Model: An Instruction Finetuned Open-Access Multilingual Language Model

2024-02-12 · arxiv-v1 · 118 页

智谱 / Z.ai
Technical Report

CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

2024-02-06 · arxiv-v3 · 21 页

DeepSeek
Technical Report

DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

2024-02-05 · arxiv-v3 · 30 页

DeepSeek
Technical Report

DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

2024-01-25 · arxiv-v2 · 23 页

DeepSeek
Technical Report

DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

2024-01-11 · arxiv-v1 · 33 页

DeepSeek
Technical Report

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

2024-01-05 · arxiv-v1 · 48 页

智谱 / Z.ai
Technical Report

CogAgent: A Visual Language Model for GUI Agents

2023-12-14 · arxiv-v3 · 27 页

Meta
Technical Report

Seamless: Multilingual Expressive and Streaming Speech Translation

2023-12-08 · arxiv-v1 · 145 页

Meta
Technical Report

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

2023-12-07 · arxiv-v1 · 15 页

Alibaba / Qwen / Wan
Technical Report

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

2023-11-14 · arxiv-v2 · 18 页

智谱 / Z.ai
Technical Report

CogVLM: Visual Expert for Pretrained Language Models

2023-11-06 · arxiv-v2 · 17 页

Google / DeepMind
Technical Report

PaLI-3 Vision Language Models: Smaller, Faster, Stronger

2023-10-13 · arxiv-v2 · 16 页

Alibaba / Qwen / Wan
Technical Report

Qwen Technical Report

2023-09-28 · arxiv-v1 · 59 页

Microsoft / Phi / MAI
Technical Report

Textbooks Are All You Need II: phi-1.5 technical report

2023-09-11 · arxiv-v1 · 16 页

Alibaba / Qwen / Wan
Technical Report

Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

2023-08-24 · arxiv-v3 · 24 页

Meta
Technical Report

Code Llama: Open Foundation Models for Code

2023-08-24 · arxiv-v3 · 48 页

Meta
Technical Report

SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

2023-08-22 · arxiv-v3 · 111 页

Google / DeepMind
Technical Report

RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control

2023-07-28 · arxiv-v1 · 26 页

Meta
Technical Report

Llama 2: Open Foundation and Fine-Tuned Chat Models

2023-07-18 · arxiv-v2 · 77 页

Microsoft / Phi / MAI
Technical Report

Textbooks Are All You Need

2023-06-20 · arxiv-v2 · 26 页

Meta
Technical Report

Simple and Controllable Music Generation

2023-06-08 · arxiv-v3 · 17 页

Meta
Technical Report

Scaling Speech Technology to 1,000+ Languages

2023-05-22 · arxiv-v1 · 41 页

Google / DeepMind
Technical Report

PaLM 2 Technical Report

2023-05-17 · arxiv-v3 · 93 页

Meta
Technical Report

DINOv2: Learning Robust Visual Features without Supervision

2023-04-14 · arxiv-v2 · 32 页

智谱 / Z.ai
Technical Report

CodeGeeX: A Pre-Trained Model for Code Generation with Multilingual Benchmarking on HumanEval-X

2023-03-30 · arxiv-v2 · 30 页

Google / DeepMind
Technical Report

PaLM-E: An Embodied Multimodal Language Model

2023-03-06 · arxiv-v1 · 18 页

Meta
Technical Report

LLaMA: Open and Efficient Foundation Language Models

2023-02-27 · arxiv-v1 · 27 页

Google / DeepMind
Technical Report

MusicLM: Generating Music From Text

2023-01-26 · arxiv-v1 · 15 页

OpenAI
Technical Report

Robust Speech Recognition via Large-Scale Weak Supervision

2022-12-06 · arxiv-v1 · 28 页

Meta
Technical Report

Galactica: A Large Language Model for Science

2022-11-16 · arxiv-v1 · 58 页

Google / DeepMind
Technical Report

Imagen Video: High Definition Video Generation with Diffusion Models

2022-10-05 · arxiv-v1 · 18 页

智谱 / Z.ai
Technical Report

GLM-130B: An Open Bilingual Pre-trained Model

2022-10-05 · arxiv-v2 · 56 页

Google / DeepMind
Technical Report

PaLI: A Jointly-Scaled Multilingual Language-Image Model

2022-09-14 · arxiv-v4 · 33 页

Meta
Technical Report

No Language Left Behind: Scaling Human-Centered Machine Translation

2022-07-11 · arxiv-v3 · 192 页

Google / DeepMind
Technical Report

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

2022-06-22 · arxiv-v1 · 49 页

智谱 / Z.ai
Technical Report

CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

2022-05-29 · arxiv-v1 · 15 页

Google / DeepMind
Technical Report

Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

2022-05-23 · arxiv-v1 · 46 页

Google / DeepMind
Technical Report

UL2: Unifying Language Learning Paradigms

2022-05-10 · arxiv-v3 · 39 页

Meta
Technical Report

OPT: Open Pre-trained Transformer Language Models

2022-05-02 · arxiv-v4 · 30 页

Google / DeepMind
Technical Report

Flamingo: a Visual Language Model for Few-Shot Learning

2022-04-29 · arxiv-v2 · 54 页

智谱 / Z.ai
Technical Report

CogView2: Faster and Better Text-to-Image Generation via Hierarchical Transformers

2022-04-28 · arxiv-v2 · 15 页

OpenAI
Technical Report

Hierarchical Text-Conditional Image Generation with CLIP Latents

2022-04-13 · arxiv-v1 · 27 页

Google / DeepMind
Technical Report

PaLM: Scaling Language Modeling with Pathways

2022-04-05 · arxiv-v5 · 87 页

Google / DeepMind
Technical Report

Training Compute-Optimal Large Language Models

2022-03-29 · arxiv-v1 · 36 页

OpenAI
Technical Report

Training language models to follow instructions with human feedback

2022-03-04 · arxiv-v1 · 68 页

Google / DeepMind
Technical Report

Competition-Level Code Generation with AlphaCode

2022-02-08 · arxiv-v1 · 74 页

Google / DeepMind
Technical Report

LaMDA: Language Models for Dialog Applications

2022-01-20 · arxiv-v3 · 47 页

Google / DeepMind
Technical Report

ShieldGemma 1 Technical Report

日期待确认 · 11 页

ByteDance / Seed
Model Card

Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity

日期待确认 · 78 页

Google / DeepMind
Technical Report

RecurrentGemma: Moving Past Transformers for Efficient Open Language Models

日期待确认 · 6 页

OpenAI
Technical Report

Improving Image Generation with Better Captions

日期待确认 · 19 页

Google / DeepMind
Technical Report

Gemini Robotics 1.5 Technical Report

日期待确认 · 62 页

Google / DeepMind
Model Card

Gemini 3.1 Flash Audio (Flash Live, TTS) Model Card

日期待确认 · 5 页

Google / DeepMind
Model Card

Gemini 2.5 Flash and Gemini 2.5 Flash Image Model Card

日期待确认 · 11 页

DeepSeek
Technical Report

DeepSeek-V3.2-Exp: Boosting Long-Context Efficiency with DeepSeek Sparse Attention

日期待确认 · 6 页