美团技术团队

美团大模型学术论文精选

Image

01. LongCat-Flash-Chat. arXiv:2509.01322.(PDF,GitHub,HuggingFace)

02.  LongCat-Flash-Thinking. arXiv:2509.18883.(PDF,GitHub,HuggingFace)

Image

01. Revisiting Scaling Laws for Language Models: The Role of Data Quality and Training Strategies. ACL 2025.(PDF)

02.  NeedleInATable: Exploring Long-Context Capability of Large Language Models towards Long-Structured Tables. NeurIPS 2025. (PDF)

03. AgentRefine: Enhancing Agent Generalization through Refinement Tuning. ICLR 2025.(PDF)

04. Earlier Tokens Contribute More: Learning Direct Preference Optimization from Temporal Decay Perspective. ICLR 2025.(PDF)

05. TODO: Enhancing LLM Alignment with Ternary Preferences. ICLR 2025.(PDF)

06. SEAS: Self-Evolving Adversarial Safety Optimization for Large Language Models. AAAI 2025.(PDF)

07. CogAtom: From Cognitive Atoms to Olympiad-level Mathematical Reasoning in Large Language Models. EMNLP2025.(PDF)

08. Mitigating Tail Narrowing in LLM Self-Improvement via Socratic-Guided Sampling. NAACL2025.(PDF)

09. SCoder: Progressive Self-Distillation for Bootstrapping Small-Scale Data Synthesizers to Empower Code LLMs. EMNLP2025 Findings. (PDF)

10. FIRE: Flexible Integration of Data Quality Ratings for Effective Pretraining. EMNLP 2025(PDF)

11. Predictor-Corrector Enhanced Transformers with Exponential Moving Average Coefficient Learning. NeurIPS 2024.(PDF)

12. Learning or Self-aligning? Rethinking Instruction Fine-tuning. ACL 2024.(PDF)

13. DolphCoder: Echo-Locating Code Large Language Models with Diverse and Multi-Objective Instruction Tuning. ACL 2024.(PDF)

Image

01. Unveiling Super Experts in Mixture-of-Experts Large Language Models. arXiv:2507.23279.(PDF)

02. EPS-MoE: Expert Pipeline Scheduler for Cost-Efficient MoE Inference. arXiv:2410.12247.(PDF)

03. FPTQ: Fine-grained Post-Training Quantization for Large Language Models. arXiv:2308.15987.(PDF)

04. A Speed Odyssey for Deployable Quantization of LLMs. arXiv:2311.09550. (PDF)

05. Flash Communication: Reducing Tensor Parallelization Bottleneck for Fast Large Language Model Inference. arXiv:2412.04964.(PDF)

06. Speculative Decoding via Early-exiting for Faster LLM Inference with Thompson Sampling Control Mechanism. ACL 2024.(PDF)

Image

01. Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation. NeurIPS 2025. (PDF)

02. InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing. arXiv:2508.14033. (PDF,HuggingFace)

03. HyperSeg: Towards Universal Visual Segmentation with Large Language Model. CVPR 2025.(PDF)

04. A Token-level Text Image Foundation Model for Document Understanding. ICCV 2025.(PDF)

05. Efficient Self-Supervised Video Hashing with Selective State Spaces. AAAI 2025. (PDF)

06. Embracing Collaboration Over Competition: Condensing Multiple Prompts for Visual In-Context Learning. CVPR 2025.(PDF)

07. Denoising with a Joint-Embedding Predictive Architecture. ICLR 2025.(PDF)

08. Enhancing Multilingual Speech Recognition Through Language Prompt Tuning and Frame-level Language Adapter. ICASSP 2024.(PDF)

09. Lumen: Unleashing Versatile Vision-centric Capabilities of Large Multimodal Models. NeurIPS 2024. (PDF)

10. MobileVLM V2: Faster and Stronger Baseline for Vision Language Model. arXiv:2402.03766v.(PDF)

11. UniViTAR: Unified Vision Transformer with Native Resolution. NeurIPS 2025.(PDF)

12. CLIP-IN: Enhancing Fine-Grained Visual Understanding in CLIP via Instruction Editing Data and Long Captions. NeurIPS 2025.(PDF)

13. Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy. NeurIPS 2025.(PDF)

14. RoboTron-Mani: All-in-One Multimodal Large Model for Robotic Manipulation. ICCV 2025.(PDF)

15. RoboTron-Drive: All-in-One Large Multimodal Model for Autonomous Driving. ICCV 2025.(PDF)

16. RoboTron-Nav: A Unified Framework for Embodied Navigation Integrating Perception, Planning, and Prediction. ICCV 2025.(PDF)

Image

01. Q-Eval-100K: Evaluating Visual Quality and Alignment Level for Text-to-Vision Content. CVPR 2025. (PDF,HuggingFace)

02. OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics. arXiv:2506.10481. (PDF,HuggingFace)

03. Ask, Fail, Repeat: Meeseeks, an Iterative Feedback Benchmark for LLMs' Multi-turn Instruction-following Ability. arXiv:2504.21625. (PDF,HuggingFace)

04. CoreCodeBench: A Configurable Multi-Scenario Repository-Level Benchmark. arXiv:2507.05281. (PDF,HuggingFace)

05. Leveraging Dual Process Theory in Language Agent Framework for Real-time Simultaneous Human-AI Collaboration. ACL 2025.(PDF)

06. Hallu-PI: Evaluating Hallucination in Multi-modal Large Language Models within Perturbed Inputs. ACM MM 2024.(PDF)

07. A Wolf in Sheep's Clothing: Generalized Nested Jailbreak Prompts can Fool Large Language Models Easily. NAACL 2024.(PDF)

Image

01. Enhancing Efficiency and Exploration in Reinforcement Learning for LLMs. EMNLP 2025. (PDF)

02. SDGO: Self-Discrimination-Guided Optimization for Consistent Safety in Large Language Models. EMNLP 2025. (PDF) 

03. When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning. EMNLP 2025. (PDF) 

04. DenoiseRotator: Enhance Pruning Robustness for LLMs via Importance Concentration. NeurIPS 2025. (PDF) 

05. AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models. ACL 2025. (PDF)

06. Don’t Half-listen: Capturing Key-part Information in Continual Instruction Tuning. ACL 2025. (PDF) 

07. PIPER: Benchmarking and Prompting Event Reasoning Boundary of LLMs via Debiasing-Distillation Enhanced Tuning. ACL 2025. (PDF)

08. CLAQ: Pushing the Limits of Low-Bit Post-Training Quantization for LLMs.(PDF)

09. A Reasoner for Real-World Event Detection: Scaling Reinforcement Learning via Adaptive Perplexity-Aware Sampling Strategy. EMNLP 2025. (PDF)