Abstract:Electroencephalography (EEG) foundation models increasingly rely on multi-dataset training and evaluation, yet public EEG datasets still lack a shared task specification layer that can turn heterogeneous recordings into reusable benchmark units. Existing standards organize files, metadata, and provenance, but they do not specify EEG tasks under a common language and rulebook, leaving critical task semantics scattered across papers, code, and manual interpretation. We investigate whether heterogeneous public EEG datasets can be standardized through a structured task specification language paired with a shared rulebook. Our methodology represents each benchmark entry as a task document synchronized with an executable task kernel, with the rulebook defining task fields, evidence requirements, document-kernel alignment, review states, and machine-checkable constraints. Using this methodology, we release a community-reviewed EEG benchmark corpus centered on 53 completed and reviewed entries with 245 task definitions spanning diverse paradigms, and we introduce NeuroDoc and NeuroAudit as the operational support layer for rulebook-guided drafting, upgrading, review, amendment, and release management. We further examine whether the resulting benchmark units can be instantiated in a shared downstream setting across four EEG foundation model backbones, providing execution-based evidence for reusable, auditable, and executable EEG benchmarking infrastructure.




Abstract:A systematic, comparative investigation into the effects of low-quality data reveals a stark spectrum of robustness across modern probabilistic models. We find that autoregressive language models, from token prediction to sequence-to-sequence tasks, are remarkably resilient (for GPT-2, test NLL increases modestly from 2.87 to 3.59 despite 50% token corruption). By contrast, under the same levels of data corruption, class-conditional diffusion models degrade catastrophically (image-label consistency plummets by 56.81% relative to baseline), while classifiers show a moderate impact that diminishes with dataset scale. To explain these discrepancies, we analyze the results through a multi-perspective lens, integrating information theory, PAC learning, and gradient dynamics. These analyses suggest that robustness is heavily influenced by two key principles: the richness of conditioning information, which constrains the learning problem, and the absolute information content of the training data, which allows the signal from correct information to dominate statistical noise.
Abstract:In this paper,we investigate a novel wireless powered mobile edge computing (MEC) system assisted by pinching antennas (PAs), where devices first harvest energy from a base station and then offload computation-intensive tasks to an MEC server. As an emerging technology, PAs utilize long dielectric waveguides embedded with multiple localized dielectric particles, which can be spatially configured through a pinching mechanism to effectively reduce large-scale propagation loss. This capability facilitates both efficient downlink energy transfer and uplink task offloading. To fully exploit these advantages, we adopt a non-orthogonal multiple access (NOMA) framework and formulate a joint optimization problem to maximize the system's computational capacity by jointly optimizing device transmit power, time allocation, PA positions in both uplink and downlink, and radiation control. To address the resulting non-convexity caused by variable coupling, we develop an alternating optimization algorithm that integrates particle swarm optimization (PSO) with successive convex approximation. Simulation results demonstrate that the proposed PA-assisted design substantially improves both energy harvesting efficiency and computational performance compared to conventional antenna systems.