Bench 2026 - 论文征稿#
Overview#
第十八届评价学:评价科学与工程大会 (Bench 2026) 将于2026年11月6日至8日在中国桂林举办。
Bench 是由国际测试委员会(BenchCouncil)主办的国际性研讨会,致力于推动评价学(Evaluatology,即评价科学与工程)的发展。本研讨会倡导跨学科采用系统化、严谨、可复现且具备科学依据的评价方法,旨在将评价学确立为一种统一的科学与工程实践范式。
Bench 2026 诚征来自广泛学科领域的原创研究与工程技术成果,涵盖计算机科学、人工智能、医学与健康、教育学、金融与经济、商业与管理、心理学、地球科学、社会科学、工程学及相关交叉领域。
本次会议特别鼓励提交与以下主题相关的跨学科研究与实践经验:评价理论与实践测试理论与实践测量理论与实践和评价、测试、测量相关的推理研究与实践求索科学体系下的评价学、计量学、测试学和推理学在各个领域的发展、应用和实践。欢迎来自所有学科相关领域的研究人员投稿。
Bench 系列会议已成功举办 十七届。往届会议论文集均发表于 Springer《计算机科学讲义》(LNCS)系列,并被 EI Compendex 收录。
往届会议论文集:
https://link.springer.com/conference/bench
全新投稿与出版模式#
Bench 2026 推出全新的 持续投稿模式。
作者全年均可随时投稿。所有稿件均将通过 BenchCouncil 投稿系统进行严格的双盲同行评审。
在 Bench 2026 评审周期内被录用的论文,在满足会议现场报告要求后,将正式纳入 Bench 2026 技术会议日程。
有关新投稿流程、出版路径、PEvaluation、TBench 特刊及报告政策的详细信息,敬请参阅:
投稿模式(新版)]()
投稿指南#
投稿语言#
所有投稿须以英文撰写。
投稿格式#
作者请依据 PEvaluation 官方稿件模板及格式规范准备文稿。
审稿政策#
Bench 2026 严格遵循双盲同行评审机制。
作者须确保所投稿件已完全匿名化。
提交的稿件中不得包含以下内容:
- 作者姓名
- 作者单位
- 致谢信息
- 基金资助声明
- 其他可能泄露身份的信息
投稿要求#
所投稿件须满足以下要求:
- 稿件须描述未在其他场合公开发表过的原创性工作。
- 投稿时,稿件不得同时处于其他会议、期刊或出版平台的审稿流程中。
- 提交版本须严格按照双盲评审要求进行匿名化处理。
- 稿件须以可打印的 PDF 文件格式提交。
- 稿件正文须包含页码。
- 图表在黑白打印条件下仍应保持清晰可读。
- 参考文献应尽可能列出全部作者,避免不必要地使用“et al.”(等)。
投稿系统#
所有投稿请通过 BenchCouncil 投稿系统提交:
https://journal.benchcouncil.org/PBench/submission
重要日期#
投稿: 全年持续接收
入选 Bench 2026 录用截止: 2026年11月1日(任意时区,AoE)
Bench 2026 会议召开时间: 2026年11月6–8日
地点: 中国桂林
征稿主题#
Bench 2026 诚征评价科学与工程领域的原创研究论文、系统论文、基准测试论文、数据集论文、测量论文、工业实践论文、综述论文、可复现性论文及立场论文。
征稿主题包括但不限于:
- Mathematical modeling and formal specification of evaluation requirements
- Development and evolution of evaluation models
- Evaluation methodology and theoretical foundations
- Design and implementation of evaluation systems
- Evaluation risk modeling and quantitative analysis
- Cost modeling and optimization for evaluations
- Accuracy modeling and error propagation analysis
- Evaluation traceability
- Identification and standardization of evaluation conditions
- Equivalent evaluation conditions and their verification
- Experimental design methodologies
- Statistical analysis techniques for evaluation
- Identification and elimination of confounding factors
- Analytical modeling and model validation
- Simulation and emulation-based modeling and validation
- Domain-specific evaluation methodologies
- Benchmark design and implementation
- Benchmark traceability
- Construction of equivalent evaluation conditions
- Evaluation metric and index system design
- Scale design and standardization
- Evaluation standard design and implementation
- Evaluation tools and toolchains
- Real-world evaluation systems
- Evaluation platforms and testbeds
- Industrial evaluation practices
- Dataset construction and development
- Dataset quality evaluation
- Dataset documentation and metadata standards
- Dataset collection, validation, and verification protocols
- Dataset reproducibility and reuse
- Dataset resampling and meta-analysis techniques
- Large-scale data generation while preserving data characteristics
- Evaluation frameworks for data-generation experiments
- Data sharing infrastructures for reproducible research
- Benchmark design and construction
- Benchmark suites and benchmark ecosystems
- Benchmark validation and maintenance
- Benchmark evolution methodologies
- Leaderboards and ranking systems
- State-of-the-art and state-of-the-practice analysis
- Industrial benchmarking practices
- Real-world application and system evaluation
- Evaluation of emerging technologies in practical scenarios
- Workload characterization
- Instrumentation, sampling, tracing, and profiling
- Measurement methodologies for large-scale systems
- Testing methodologies and frameworks
- Measurement-driven knowledge discovery
- Performance modeling and bottleneck analysis
- Scalability and efficiency evaluation
- Monitoring and visualization of measurement data
- Reproducible measurement practices
- Re-evaluation of previous empirical measurements and conclusions
- Benchmark-driven algorithm evaluation
- Standardized evaluation of algorithms
- Algorithm performance analysis
- Generalization, robustness, fairness, and reliability evaluation
- Hardware-aware algorithm evaluation and co-design
- Accuracy-cost and efficiency-quality trade-off analysis
- Adaptive evaluation under dynamic environments
- Reproducibility of algorithm evaluation
- Optimization driven by evaluation feedback
- Foundation model evaluation
- Large language model evaluation
- AI benchmark development
- Agent evaluation
- Multimodal AI evaluation
- AI safety evaluation
- Robustness evaluation
- Fairness evaluation
- Explainability evaluation
- Human-centered evaluation
- Real-world AI system evaluation
Bench 2026 welcomes evaluation research and engineering practices in, but not limited to:
- Computer science
- Artificial intelligence
- Medicine and healthcare
- Biology
- Education
- Finance and economics
- Business and management
- Psychology
- Social sciences
- Earth sciences
- Environmental sciences
- Transportation
- Energy
- Manufacturing
- Digital humanities
- Other scientific and engineering disciplines
联系方式#
投稿相关咨询,请联系: yangzhengxin@ict.ac.cn
