Back to the news portal
Climate & EnergyResearch paperResearchSource analysisNorth AmericaUnited States

Can AI control survive outside simulation for wave energy?

A US wave-flume study exposed a deep-reinforcement-learning controller's first failures, retrained it for physical losses and recovered competitive mechanical power. It tested one scaled device and one irregular sea state—not ocean deployment or electricity output.

By The Impact of AI Research DeskReleased 2 October 2026 at 23:15 BST7 min read3 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesWave energyReinforcement learningSim-to-realRenewable energyControl systems

Research topic

Whether a deep-reinforcement-learning power-take-off controller remains robust and productive when moved from simulation to a physical wave-energy converter

The Impact of AI research cover asking whether AI control can survive outside simulation, with a conceptual buoy, simulated waves and a physical wave tank.
AI-generated editorial illustration. The buoy, simulated waves and tank are conceptual and do not depict the tested apparatus, an ocean deployment or measured performance.

At a glance

  • 1The initial controller often failed in Oregon State University's wave flume because the simulation omitted drivetrain losses, random starting conditions, nonlinear events and process noise.
  • 2After training was widened, the controller matched an experimentally tuned feedback benchmark on most of six regular conditions, produced 56.8% more mechanical power in the most energetic one, and beat it in three runs of one irregular sea state.
  • 3This was a scaled, heave-only tank experiment measuring mechanical power—not an ocean trial, grid electricity result, durability test or economic assessment.

Living evidence record

Impact record IAI-1R6OY4D

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

2 October 2026

Source trail

3 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

The study tested the gap between a model and moving water

Wave-energy converters can extract more power when their power-take-off system adjusts to changing motion without pushing the device beyond physical limits. Deep reinforcement learning is attractive because it can learn control actions through repeated interaction in a simulator. The hard question is whether a policy trained in a clean numerical environment still works when friction, sensor noise, irregular waves and uncertain starting states enter the system.

Researchers tested that question on the open-source Laboratory Upgrade Point Absorber, or LUPA, at Oregon State University's O.H. Hinsdale Wave Research Laboratory. The device was configured as a single body moving only in heave. A deep Q-network adjusted the gains of a time-varying proportional-integral controller for the power-take-off. Training took place offline in MATLAB and Simulink; execution ran in real time on a Speedgoat machine.[1][2]

Six regular conditions and one irregular state formed the test

The campaign used six regular wave combinations: heights of 0.15 or 0.20 metres paired with selected periods from 1.75 to 2.75 seconds. It also used one irregular Pierson–Moskowitz sea state with a significant height of 0.21 metres and peak period of 3.09 seconds. Conditions were scaled from the operating range expected for LUPA at PacWave South off Newport, Oregon; they were not full-scale ocean waves.

Regular-wave runs lasted 600 seconds and irregular-wave runs 800 seconds. Each test was repeated three times. The benchmark was an experimentally tuned feedback controller using the same physical setup and control structure, with gains swept through tank tests. That is a stronger practical comparator than an idealised theoretical controller because both systems experienced the same drivetrain losses and stroke constraints.[2]

The first transfer from simulation largely failed

In numerical training, the deep-RL controller appeared close to a linear impedance-matching reference while respecting the device's motion limit. For one regular condition, a 0.20-metre wave and 2.25-second period, simulation predicted average mechanical power of 92.66 watts over the final 100 seconds. That result looked promising but described a model, not the flume.

When pretrained agents were deployed directly, most conditions ended in control failure, defined as both controller gains approaching their imposed limits; without those bounds the system could become unstable. In one comparison, simulation predicted 51.4 watts while the experiment produced 14.9 watts. The tuned benchmark also produced 14.9 watts where an idealised model predicted 55.4 watts. The transfer error exposed substantial missing physical behaviour.[2]

Training was redesigned around uncertainty

The team traced the mismatch to nonlinear power-take-off loss, random starting conditions, nonlinear events and process noise. Belt and motor losses reduced motion much more than the original model expected, and varied with the wave and control policy. In two examples, the model overpredicted peak-to-peak uncontrolled heave by roughly 80%; under controlled motion the mismatch exceeded 180%. Researchers inserted an empirical range of extra damping, explicitly describing it as an ad-hoc approximation rather than a complete drivetrain model.

The revised simulator randomised damping, controller activation time and initial conditions. It injected short disturbances and added normally distributed process noise. The paper says the disturbance heuristic was not physically consistent; it merely taught recovery from patterns seen in the flume. The reward also changed from instantaneous power to power averaged over a sliding window, reducing incentives to chase spikes. Eight observation-state combinations were tested, and different ones worked best for regular and irregular waves.[2]

The revised controller recovered competitive results

After retraining, the deep-RL controller produced mechanical power close to the tuned feedback controller across most regular conditions. In the most energetic regular condition, with a 0.20-metre height and 2.25-second period, it produced about 56.8% more power while using proportional and integral control within physical constraints. The benchmark used damping only there because integral action risked excessive motion.

For the single irregular sea state, three random-seed runs produced approximately 11.6, 14.0 and 15.5 watts of average mechanical power. The feedback benchmark produced 5.4 watts. Those repetitions show consistency within one condition, not general performance across real seas. The authors call the irregular-wave validation preliminary. Real-time computation itself was not the bottleneck: reported processing time was 0.008 milliseconds within a one-millisecond sample interval.[2]

Mechanical power is not delivered electricity

The main reward and comparison concerned mechanical power. The setup did not provide dedicated electrical-power measurements, and the paper separately explored a force penalty as a proxy for drivetrain loss. In one repeated observation, small controller-gain oscillations increased mechanical power by 26%, but the authors warn that the behaviour could worsen electrical efficiency, fatigue or constraint violations.

The study therefore did not show a 56.8% increase in electricity delivered to a grid, lower lifetime cost or higher commercial capacity factor. A controller that extracts more short-run mechanical power can still lose energy in conversion or shorten equipment life. Wave-energy projects are judged by reliability, maintenance, survivability and cost as well as capture in a controlled tank.[2]

A useful validation method, not an ocean breakthrough

The strongest contribution is procedural: publish the initial sim-to-real failure, identify missing variability, retrain against uncertainty and retest on the same apparatus. Many reinforcement-learning control papers stop at simulation, where an agent can appear optimal because the environment shares its training assumptions. The flume showed how quickly that performance could disappear.

But the experiment used one scaled device, one degree of freedom and a limited set of laboratory waves. It did not expose the controller to storms, currents, biofouling, sensor drift, mooring dynamics or long-duration degradation. The US Department of Energy, Pacific Ocean Energy Trust and TEAMER programme funded the work; authors declared no competing interests and deposited data in the Marine and Hydrokinetic Data Repository.[2][3]

What would change the assessment

Confidence would rise with many irregular sea states, a physically grounded drivetrain model and direct electrical-power measurement. Experiments should report control violations, loads, fatigue proxies and recovery from sensor or actuator faults. Comparisons with model-predictive, adaptive and robust controllers would show whether reinforcement learning adds value beyond the selected feedback baseline.

The decisive evidence would be a longer prototype-scale field trial with a locked controller facing unseen ocean conditions, measuring net electricity, availability, component wear, maintenance and economics. Until then, the paper shows that a controller can be repaired after a revealing laboratory failure and beat one practical benchmark under selected conditions—not that AI has solved wave energy.[2]

What this means for people

  • Better control could improve usable output from wave-energy devices, but this study does not establish cheaper or more reliable electricity.
  • Publishing the failed first transfer gives engineers a clearer account of safety limits and hidden losses than simulation-only results.
  • Communities and developers should not treat tank-scale mechanical power as evidence of an ocean-ready commercial system.

Global context

The work was funded and tested in the United States using conditions scaled from an Oregon site. Wave resources, ports, supply chains and permitting differ across regions. The sim-to-real lesson is globally relevant, but deployment anywhere would require device- and site-specific control, safety and economic validation.

What the evidence does not yet show

  • The experiment used one scaled LUPA configuration in heave-only mode and only one irregular sea state.
  • Results concern short wave-flume runs and mechanical power, not net electricity, ocean durability, availability or cost.
  • Several training disturbances and drivetrain-loss ranges were empirical heuristics rather than physically complete models.
  • The controller was compared with one experimentally tuned feedback design; other modern strategies were not tested head-to-head.

What to watch next

  • Direct electrical-power measurements and a validated nonlinear power-take-off model.
  • A broader matrix of irregular waves, faults, currents and multi-axis motion.
  • Long-duration prototype-scale ocean trials with reliability, fatigue and maintenance outcomes.
  • Independent comparisons against model-predictive, adaptive and robust control systems.

Evidence trail

Sources used for this report

Links checked 2 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.