• Xeuron logo
v5-202608201923
Feed
  • Home
  • Papers
  • Discussions
  • Leaderboard
  • Workplace
Discover
  • Ask-XeuronAI
  • Agent Library
  • Popular
  • Hot & Trending
  • Explore
  • My Extractions
Admin
  • Dashboard
  • Metrics
  • Cron Jobs
  • Model Registry
Create
  • SubXeurons
    • iPSC-Cardio Cells
    • HALO: A Unified Visio
  • Publications
    • vision foundation model for single-cell biology via spatial gene cartography
    • Adoption and Use of LLMs at an Academic Medical Center
    • Toward AI-Driven Digital Organism
    • You Can Run, You Can Hide: The Epidemiology and Statistical Mechanics of Zombies
    • embryonic stem cell-derived cardiac organoids via synthetic guidance
    • In vitro generation of human pluripotent stem cell derived lung organoids
    • Generating Self-Assembling Human Heart Organoids Derived from Pluripotent Stem Cells
    • SMAD4: A Critical Regulator of Cardiac Neural Crest Cell Fate and Vascular Smooth Muscle Differentiation. bioRxiv
    • Insights into AI Agent Security from a Large-Scale Red-Teaming Competition
    • TxPert: using multiple knowledge graphs for prediction of transcriptomic perturbation effects
    • Self-organizing human heart assembloids with autologous and developmentally relevant cardiac neural crest-derived tissues
    • Path Planning of Cleaning Robot with Reinforcement Learning
    • Reinforcement Learning Approaches in Social Robotics
    • Robotic Packaging Optimization with Reinforcement Learning
    • A Concise Introduction to Reinforcement Learning in Robotics
    • Robot-R1: Reinforcement Learning for Enhanced Embodied Reasoning in Robotics
    • Robotic Surgery With Lean Reinforcement Learning
    • Residual Reinforcement Learning for Robot Control
    • Autonomous robotic nanofabrication with reinforcement learning
    • Heterogeneous Multi-Robot Reinforcement Learning
    • Robot Air Hockey: A Manipulation Testbed for Robot Learning with Reinforcement Learning
    • Reinforcement learning for freeform robot design
    • Geometric Reinforcement Learning For Robotic Manipulation
    • On-Robot Bayesian Reinforcement Learning for POMDPs
    • Efficient Content-Based Sparse Attention with Routing Transformers
    • A foundation model of transcription across human cell types
    • Transformer AI
    • HALO, a unified VLA model that enables embodied multimodal chain-of-thought (EM-CoT) reasoning through a sequential process of textual task reasoning, visual subgoal prediction for fine-grained guidan
    • HALO: A Unified Vision-Language-Action Model for Embodied Multimodal Chain-of-Thought Reasoning
  • Events
    • No events yet
HomeSearchEventsProfileCreate
  • Home
  • Papers
  • Discussions
  • Leaderboard
  • Workplace
  • Robotic Packaging Optimization with Reinforcement Learning

    Preprint[2023]
    ·Source·PDF|

    AI Summary

    Intelligent manufacturing is becoming increasingly important due to the growing demand for maximizing productivity and flexibility while minimizing waste and lead times. This work investigates automated secondary robotic food packaging solutions that transfer food products from the conveyor belt into containers. A major problem in these solutions is varying product supply which can cause drastic productivity drops. Conventional rule-based approaches, used to address this issue, are often inadequate, leading to violation of the industry's requirements. Reinforcement learning, on the other hand, has the potential of solving this problem by learning responsive and predictive policy, based on experience. However, it is challenging to utilize it in highly complex control schemes. In this paper, we propose a reinforcement learning framework, designed to optimize the conveyor belt speed while minimizing interference with the rest of the control system. When tested on real-world data, the framework exceeds the performance requirements (99.8% packed products) and maintains quality (100% filled boxes). Compared to the existing solution, our proposed framework improves productivity, has smoother control, and reduces computation time.

    AI Metadata Extraction

    Extract authors, key findings, references, and an executive summary using AI.

    Preparing…Generating PDF…
    Version:· 1 version extracted
    Extraction v1google/gemini-3.1-flash-lite6/27/2026

    Executive Summary

    This paper investigates the feasibility of applying Reinforcement Learning (RL) to solve complex robotic packaging problems, specifically optimizing conveyor belt speeds under varying product inflow conditions. The authors propose an RL framework designed to integrate seamlessly into existing industrial control schemes by utilizing a 'planned delay' mechanism. This approach minimizes interference with the machine's existing internal scheduling, allowing the agent to provide control signals that are compatible with the real-time requirements of the packaging line. To ensure the system works in realistic environments, the researchers trained their RL agent using simulated randomized scenarios that replicate real-world data distributions. Validation was conducted using actual operational data from a food packaging machine, comparing the RL solution against an industry-standard rule-based baseline. The results demonstrate that the RL approach consistently outperforms the baseline, achieving a 99.94% performance rate (exceeding the 99.8% requirement) and reducing product loss by over 93%. Additionally, the RL framework provides smoother mechanical control and reduced computation times compared to the traditional approach. The findings indicate that reinforcement learning can effectively handle the high-dimensional, time-sensitive nature of industrial robotics when combined with careful state representation, reward design, and delay management. By avoiding the myopic, reactive nature of simple rule-based systems, the RL agent proactively adjusts belt speeds to prevent product bottlenecks. This study represents a significant step toward the practical deployment of intelligent, data-driven controllers in competitive manufacturing environments.

    Accurate?

    Authors (5)

    Eveline DrijverFirst Author

    Cognitive Robotics Department, Delft University of Technology, 2628 CD Delft, The Netherlands

    E.A.Drijver@gmail.com

    Rodrigo Pérez-Dattari

    Cognitive Robotics Department, Delft University of Technology, 2628 CD Delft, The Netherlands

    Jens Kober

    Cognitive Robotics Department, Delft University of Technology, 2628 CD Delft, The Netherlands

    Accurate?

    Abstract

    Intelligent manufacturing is becoming increasingly important due to the growing demand for maximizing productivity and flexibility while minimizing waste and lead times. This work investigates automated secondary robotic food packaging solutions that transfer food products from the conveyor belt into containers. A major problem in these solutions is varying product supply which can cause drastic productivity drops. Conventional rule-based approaches, used to address this issue, are often inadequate, leading to violation of the industry’s requirements. Reinforcement learning, on the other hand, has the potential of solving this problem by learning responsive and predictive policy, based on experience. However, it is challenging to utilize it in highly complex control schemes. In this paper, we propose a reinforcement learning framework, designed to optimize the conveyor belt speed while minimizing interference with the rest of the control system. When tested on real-world data, the framework exceeds the performance requirements (99.8% packed products) and maintains quality (100% filled boxes). Compared to the existing solution, our proposed framework improves productivity, has smoother control, and reduces computation time.

    Accurate?

    Fields of Study

    Robotic Packaging OptimizationReinforcement Learning for ControlIndustrial RoboticsAutomated ManufacturingControl Systems EngineeringRoboticsArtificial Intelligence
    Accurate?

    Key Findings (20)

    1.The proposed RL framework achieved a 99.94% performance rate, exceeding the 99.8% industry requirement.

    2.The RL solution achieved a 100% quality rate, ensuring no empty or partly filled boxes left the machine.

    3.The RL method resulted in a 93.26% reduction in lost products compared to the rule-based baseline.

    Accurate?

    Discussion & Future Directions

    The paper concludes that the proposed reinforcement learning framework effectively manages control delays and sparse rewards in the interdependent control schemes typical of robotic packaging. The learned policy provides predictive behavior under varying product supply while ensuring industrial safety and quality requirements are met. Future work will focus on transferring this learned policy from simulation to physical robotic platforms, addressing the gap between simulated training and real-world deployment through further data integration and handling of novel edge cases.

    Accurate?

    References (23)

    1. [1]Brunke, L., Greeff, M., Hall, A. W., Yuan, Z., Zhou, S., Panerati, J., and Schoellig, A. P. (2022). Safe learning in robotics: From learning-based control to safe reinforcement learning. Annual Review of Control, Robotics, and Autonomous Systems, 5, 411–444.
      Create publication
    2. [2]Bouteiller, Y., Ramstedt, S., Beltrame, G., Pal, C., and Binas, J. (2021). Reinforcement learning with random delays. In Int. Conf. Learning Representations (ICLR).
      Create publication
    3. [3]Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y. (2018). Safe exploration in continuous action spaces. arXiv preprint arXiv:1801.08757.
      Create publication
    Accurate?

    Sections

    Executive SummaryAuthorsAbstractFields of StudyKey FindingsDiscussionReferences