Goto

Collaborating Authors

 Large Language Model



Prism: A Framework for Decoupling and Assessing the Capabilities of VLMs

Neural Information Processing Systems

Limitations in either capability can impede the overall performance of a VLM. A systematic evaluation of the perception and reasoning capabilities is crucial to provide valuable insights for future model optimization.






Toward a Stable, Fair, and Comprehensive Evaluation

Neural Information Processing Systems

Overcoming this challenge, existing object hallucination evaluation methods average the results obtained from a set of instructions. However, these methods fail to provide consistent evaluation across instruction sets that generate image descriptions of significantly different lengths.


FlexCap: Describe Anything in Images in Controllable Detail

Neural Information Processing Systems

We demonstrate FlexCap's effectiveness in several applications: first, it achieves strong performance in dense captioning tasks on the Visual Genome dataset. Second, we show how FlexCap's localized


FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models Supplementary Materials 1 Dataset 1.1 Links and Preservation

Neural Information Processing Systems

The croissant metadata record is available at croissant. We chose GitHub and Google Drive respectively to store our code and dataset. Both are widely recognized as reliable data storage platforms, ensuring long-term preservation. We highly recommend downloading the raw data directly and following the provided instructions to simplify the data processing steps. Our dataset is structured as follows: the local directory contains client-specific data for local training, while all clients aggregates data from all clients for federated learning.


FedLLM-Bench: Realistic Benchmarks for Federated Learning of Large Language Models Rui Ye1 Rui Ge

Neural Information Processing Systems

Based on FedLLM-Bench, we conduct experiments on all datasets to benchmark existing FL methods and provide empirical insights (e.g., multilingual collaboration). We believe that our FedLLM-Bench can benefit the FedLLM community by reducing required efforts, providing a practical testbed, and promoting fair comparisons.