Geri Dön

Hava tahmin modelleri ve yüzdelik dilim regresyonu temelli hibrit güneş enerjisi öngörü sistemi geliştirilmesi ve uygulamaları

Development and applications of a hybrid solar energy forcasting system integrating weather models and quantile regression

  1. Tez No: 1020044
  2. Yazar: AYŞEGÜL BİÇER
  3. Danışmanlar: PROF. DR. AHMET DURAN ŞAHİN
  4. Tez Türü: Yüksek Lisans
  5. Konular: Meteoroloji, Enerji, Meteorology, Energy
  6. Anahtar Kelimeler: Belirtilmemiş.
  7. Yıl: 2026
  8. Dil: Türkçe
  9. Üniversite: İstanbul Teknik Üniversitesi
  10. Enstitü: Lisansüstü Eğitim Enstitüsü
  11. Ana Bilim Dalı: Meteoroloji Mühendisliği Ana Bilim Dalı
  12. Bilim Dalı: Atmosfer Bilimleri Bilim Dalı
  13. Sayfa Sayısı: Belirtilmemiş.

Özet

Yenilenebilir enerji kaynakları, iklim değişiminden en az şekilde etkilenmek için kilit görevi görmektedir. Günümüzde artan teknoloji sayesinde yenilenebilir enerji kaynaklarına erişim kolaylaşmıştır. Ülkemizde hem sanayi hem de kişisel amaçlar için en çok kullanılan yenilenebilir enerji kaynağı güneş enerjisidir. Güneş enerjisinden üretilen elektriğin tahminlenmesi ise şebeke güvenliği ve ekonomik nedenler dolayısıyla önem arz etmektedir. Günümüzde, tahminleme için en çok kullanılan yöntem makine öğrenimi ve derin öğrenme metotlarıdır. Tezin ana amacı, güneş enerjisi santrallerinin gün öncesi elektrik üretimlerini tahmin etmek için kullanılan nokta tahminlerinin yetersiz kaldığı durumları belirlemek ve olasılık ve hava durumuna bağlı karar mekanizmasını geliştirmektir. Bu çalışmada, 3 adet güneş enerjisi santrali (GES) kullanılmış olup bunlar Konya (18MWe), Sivas (9MWe) ve Antalya (5,6MWe) illerinde yer almaktadır. Veri kaynağı olarak Meteoroloji Genel Müdürlüğü (MGM)'nden alınan gözlem verileri ile Avrupa Orta Vadeli Hava Tahminleri Merkezi (ECMWF)'ne ait Entegre Tahmin Sistemi (IFS) modelinin yüksek çözünürlüklü (HRES) geçmişe yönelik çıktıları kullanılmıştır. Kullanılan iki veri setide 2 yıllık (2023 Ekim – 2025 Ekim) ve saatlik çözünürlüktedir. MGM ve IFS HRES veri setleri için berrak hava indeksi, güneş yükseklik açısı, son 3 saatteki bulutluluk miktarının standart sapması ile tarih ve saatin trigonometrik dönüşümleri özellik mühendisliği olarak eklenmiştir. Her iki veri seti, ham veri ve Çekirdek Temel Bileşen Analizi (ÇTBA) ile boyut indirgenerek elde edilen şekilde 2 farklı yolla modellere girdi olarak eklenmiştir. Nokta tahminlerinde doğrusal model olarak çoklu lineer regresyon (LR), doğrusal olmayan model olarak Ekstrem Gradyan Güçlendirme (XGBoost) modeli kullanılmıştır. Test periyodu ve 160 günlük gün öncesi operasyonel tahminler için Bulanık C Ortalamaları (BCO) ile ayrılan kümeler bazında hata analizi yapılmıştır. Noktasal modellerin öğrenme eğrileri incelenmiş ve lineer modellerin düzgün öğrenememe, XGBoost modellerinin ise ezberleme yaptığı görülmüştür. Olasılık uzayı elde etmek için Yüzdelik Dilim Regresyonu (YDR) kullanılmıştır. Doğrusal olarak, Çoklu Doğrusal Yüzdelik Dilim Regresyonu (ÇDYDR); doğrusal olmayan model olarak XGBoost tabanlı Yüzdelik Dilim Regresyonu (XGB YDR) eğitilmiştir. Bu eğitimler 10 ve 90 arasındaki 10'ar artışla tüm yüzdelik dilimlerin tahminlenmesi için tekrarlanmıştır. Sürekli Dereceli Olasılık Skoru (CRPS), Tahminlenen Aralık Kapsama Olasılığı (PICP) ve Tahmin Aralığı Normalize Edilmiş Ortalama Genişlik (PINAW) metrikleri ile performansları ölçülmüştür. Hava rejimleri, BCO kümeleme yöntemi kullanılarak; açık hava, bulutlu hava, geçişli bulutlu hava, karlı hava ve düşük ışınım havası olmak üzere 4 veya 5 kümeye ayrılmıştır. Her rejim için ortalama mutlak hata (MAE)'yı minimize eden yüzdelik dilim belirlenerek bir karar mekanizması oluşturulmuştur. Bu karar verme mekanizması ile 160 günlük operasyonel gün öncesi tahminler yapılmış ve nokta tahmini ile karar mekanizması tahminleri kıyaslanmıştır. Çalışmadan elde edilen sonuçlara göre; nokta tahmininde genellikle XGBoost modelleri LR'den daha iyi sonuçlar gösterirken ÇTBA girdisi çoğunlukla ham veriden daha kötü sonuçlar sağlamıştır. Nokta tahminlerinde 3 GES arasından en iyi normalize ortalama mutlak Hata (NMAE) %5,2 ile Hamal GES IFS HRES ham veri girdili XGBoost modelinden elde edilmiştir. Alibeyhüyüğü GES içinde en iyi model IFS HRES ham veri girdili XGBoost modeli olmuştur. Serra GES, dağlık ve karma iklimi nedeniyle eğitim ve operasyonel süreçlerde hem nokta tahminlerinde hem de karar mekanizması tahminlerinde en zorlu GES olmuştur. Noktasal tahmin modellerinde diğer 2 GES'den ayrışarak en iyi sonucu IFS HRES ÇTBA girdili LR modelinden almıştır. ÇTBA girdisi genellikle nokta tahminlerinin performansını düşürürken yüzdelik dilimler için lineer kalibrasyona yardımcı görev üstlenebildiği görülmüştür. Yüzdelik dilim regresyonları içinde XGBoost tabanlı modeller kendisine aşırı güvenen yüzdelik dilimler tahmin etmiş, çoklu doğrusal modeller ise geniş tahmin aralıkları sunarak temkinli davranmıştır. BCO ve YDR sayesinde elde edilen karar mekanizması sonuçları ile genel olarak nokta tahmini performansları iyileştirilmiştir. En belirgin iyileşmeler, ÇTBA girdili modellerde ve MGM veri setinde görülmüştür. En yüksek MAE iyileşme oranı %17,8 ile Hamal GES'te ÇTBA girdili ÇDYDR karar modeli ile sağlanmıştır. IFS HRES ve XGBoost ham veri modelleri gibi nokta tahminini güçlü yapan modellerde iyileşme oranı daha az olarak görülmüştür. Serra GES için ise genel olarak deterministik nokta tahminleri daha iyi performans sergilemiştir. Modeller birbirleri ile karşılaştırılabilmesi ve davranış kalıplarının görülebilmesi için optimizasyon yapılmadan taban model olarak kullanılmıştır. Gelecek çalışmalarda taban modelleri optimizasyon yapılarak geliştirilecektir. Serra GES'te gözlenen zorluklardan yola çıkarak gelecek çalışmalarda dağlık ve karma iklim bölgelerine özgü özellik mühendisliği ve veri temsili stratejileri araştırılacaktır. Ayrıca, sayısal hava tahmini yöntemlerinin yanı sıra uydu kaynaklı radyasyon tahmin ürünlerinin de veri seti genişletme amacıyla sisteme dâhil edilmesi planlanmaktadır. Son olarak, çalışamda kullanılan GES sayısının artırılması ve farklı kapasite ile teknolojilere sahip santrallerin analize dâhil edilmesi, önerilen yöntemin istatistiksel güvenilirliğini ve genellenebilirliğini güçlendirecektir.

Özet (Çeviri)

Renewable energy sources are among the most critical energy alternatives of our time, offering a means to reduce dependence on fossil fuels and mitigate the adverse effects of climate change. Solar energy, in particular, represents an exceptionally suitable resource for Türkiye given its geographical location and climatic characteristics. Thanks to advancing technology and declining system costs, the installed capacity of solar plants (SPP) has been growing rapidly year after year. This growth makes accurate and reliable forecasting of electricity generation from SPPs both technically and economically essential. Solar energy generation forecasting is vital for grid operators seeking to minimize balancing costs, energy market participants formulating accurate bidding strategies, and plant managers conducting maintenance planning. Inaccurate forecasts not only jeopardize grid security but also drive up imbalance costs. For this reason, the literature applies a wide range of linear and non linear machine learning methods – combined with various data sources and meteorological clustering strategies – to the solar energy forecasting problem. Point forecasting models predict a single generation value and therefore can not capture uncertainty; their performance degrades notably under transitional weather conditions such as overcast, snowy, or rainy weathers. To address this shortcoming, probabilistic forecasting methods have attracted growing interest. Rather than providing a single value, probabilistic forecasting presents decision – makers with the full spectrum of possible generation scenarios, enabling more informed and risk – sensitive decisions. The core motivation of this study is to systematically identify weather conditions under which point forecasting models fall short, model these conditions within a probabilistic framework, and select the optimal forecast value through a decision - theory – based mechanism. Three solar power plants (SPP) were selected to represent different climatic regions of Türkiye; the 18 MWe Alibeyhüyüğü SPP in Konya, the 9 MWe Hamal SPP in Sivas, and the 5.6 MWe Serra SPP in Antalya. The distinct geographical and climatic conditions of these three plants were a deliberate choice for testing the generalizability of the proposed method. Two different data sources were used: hourly observational data from the Turkish State Meteorological Service (MGM) and high – resolution (HRES) reanalysis outputs from the European Centre of Medium – Range Weather Forecasts' (ECMWF) Integrated Forecast System (IFS) model. Both datasets cover a two-year period from October 2023 to October 2025 at hourly resolution. The ECMWF IFS HRES dataset includes global radiation, 2-meter temperature, relative humidity, total cloud cover, 10-meter wind speed, total precipitation, total snowfall, and surface pressure. The MGM dataset contains hourly global radiation, current snow depth, actual pressure, relative humidity, wind speed, temperature, total precipitation, and cloud cover (in oktas). Both datasets were augmented through feature engineering with a clear-sky index, solar elevation angle, the standard deviation of cloud cover over the preceding three hours, and trigonometric date-time transformations. Two input formats were prepared for the models: raw data and dimensionality-reduced data via Kernel Principal Component Analysis (KPCA). KPCA aims to reduce input dimensionality while preserving information within nonlinear data structures. For point forecasting, multiple linear regression (LR) was used as the linear model and XGBoost as the nonlinear model. For probabilistic forecasting, multiple linear quantile regression (MLQR) and XGBoost-based quantile regression were trained across all quantiles from the 10th to the 90th percentile in steps of ten. Probabilistic model performance was evaluated using CRPS, PICP, and PINAW metrics. Fuzzy C-Means (FCM) clustering was employed to identify weather regimes, classifying conditions into four to five clusters: clear-sky, overcast, transitionally cloudy, snowy, and low-irradiance weather. A decision mechanism was constructed by identifying the quantile that minimizes MAE for each regime, and this mechanism was applied to 160 days (1 November 2025 – 9 April 2026) of operational day-ahead forecasts. Examining the point forecasting results, XGBoost models generally outperformed linear regression models. However, dimensionality reduction via KPCA yielded lower point forecasting performance than raw data in most cases — a finding suggesting that dimensionality reduction can diminish the direct information content available for point forecasting. Among the three plants, the best point forecasting performance was achieved at Hamal SPP with an NMAE of 5.2%, obtained from the XGBoost model with IFS HRES raw data input. The best-performing model for Alibeyhüyüğü SPP was likewise the XGBoost model with IFS HRES raw data input. Serra SPP proved the most challenging plant in both training and operational phases, owing to its mountainous terrain and mixed climate. Unlike the other two plants, Serra SPP's best point forecasting result came from a linear regression model with IFS HRES KPCA input — indicating that dimensionality reduction can exert a regularizing effect on linear models under complex and variable weather conditions. Analysis of the learning curves showed that linear models suffered from underfitting, while XGBoost models exhibited a tendency toward overfitting. In evaluating the probabilistic forecasting models, XGBoost-based models were found to produce overconfident quantiles, whereas multiple linear quantile regression offered a more conservative approach through wider prediction intervals. While KPCA input had a detrimental effect on point forecasts, it was observed to potentially improve calibration quality in linear quantile regression models. The FCM clustering and quantile regression-based decision mechanism generally improved upon point forecasting performance, with the most pronounced gains observed in KPCA-based models and the MGM dataset. The highest MAE improvement rate — 17.8% — was achieved at Hamal SPP through the MLQR decision model with KPCA input. For models already delivering strong point forecasts, such as the IFS HRES and XGBoost raw data combination, the improvement brought by the decision mechanism was comparatively limited. For Serra SPP specifically, direct point forecasting models outperformed the decision mechanism approach. This finding reveals that the proposed method can behave differently depending on geographical and climatic conditions. Additionally, the use of observational data compiled from multiple MGM stations was found to reduce representational capacity over the plant site, particularly exacerbating forecasting difficulties in Serra SPP's complex geography. All models employed in this study were used without any hyperparameter optimization, in order to keep behavioral patterns comparable and establish a baseline. Future work will include systematic optimization of both point and probabilistic forecasting models, with the goal of obtaining results closer to the decision mechanism's potential performance ceiling. Drawing on the challenges observed at Serra SPP, future research will also explore feature engineering and data representation strategies tailored to mountainous and mixed-climate regions. In cases where relying on data from a single meteorological station nearest to the plant proves insufficient, multi-station integration or distance-weighted interpolation methods may be considered. Furthermore, satellite-derived irradiance forecast products are planned to be incorporated alongside high-resolution NWP data to expand the dataset. Finally, increasing the number of SPPs beyond the three used in this study and incorporating plants with varying capacities and technologies would substantially strengthen the statistical robustness and generalizability of the proposed method. From a practical standpoint, the proposed decision mechanism offers tangible value to grid operators and energy market participants operating under Türkiye's day- ahead and intraday electricity markets. By systematically flagging the weather regimes in which point forecasts are least reliable and substituting a probabilistically informed estimate in those regimes, the method can help plant operators reduce imbalance penalties without requiring a wholesale replacement of their existing point – forecasting infrastructure. This selective intervention strategy – applying the decision mechanism only where it adds value, rather than uniformly across all conditions – distinguishes the present approach from prior studies that adopt either a purely deterministic or a purely probabilistic forecasting pipeline. The cross – plant comparision further underscores that no single combination of data source, dimensionality – reduction strategy, and model type is universally optimal; rather, the best configuration is contingent on a plant's geographical and climatic profile. This reinforces the broader literature's emphasis on site – specific model selection and cautions against generalizing forecasting recommendations derived from a single location to plants situated in markedly different terrain. Methodologically, the study contributes a reproducible framework that couples unsupervised weather – regime clustering with quantile – based decision rules, offering a template that can be extended to other renewable sources, such as wind, where similar weather – driven forecast degradation is observed. Taken together, these findings provide both a practical decision – support tool for operational use and a methodological foundation for future probabilistic forecasting research in geographically diverse solar energy contexts. Overall, this thesis demonstrates that combining physics – informed feature engineering, unsupervised weather clustering, and quantile – based decision rules can meaningfully narrow the gap between theoretical forecasting accuracy and operational reliability. The comparative analysis across three climatically distinct power plants shows that forecasting performance cannot be optimized through a one – size – fits – all configuration, but instead requires site – aware model and input selection. By validating the proposed framework against 160 days of real operational forecasts rather than relying solely on offline test – set performance, this study offers evidence that is directly transferable to day – ahead market operations. These contributions collectively support more resilient and economically efficient integration of solar energy into Türkiye's power grid.

Benzer Tezler

  1. Katı yakıtların akışkan yatakta yakılması ve katı miktarının yatak yüksekliğine etkisi

    Combustion of solid fuels in the fkuidized beds and the effects of solid quantity to the height of bed

    KAMİL BEKİR KOÇ

    Yüksek Lisans

    Türkçe

    Türkçe

    1987

    Makine MühendisliğiGazi Üniversitesi

    Makine Eğitimi Ana Bilim Dalı

    YRD. DOÇ. DR. ALİ YÜCEL UYAREL

  2. Thlaspi jaubertii hedge üzerinde morfolojik araştırmalar

    Başlık çevirisi yok

    O.CEM ERGÜL

    Yüksek Lisans

    Türkçe

    Türkçe

    1987

    BiyolojiUludağ Üniversitesi

    Biyoloji Ana Bilim Dalı

    DOÇ. DR. ALİ ÇIRPICI

  3. Statik elektrikle yüklenmiş sıvı ilacın pamuk ilaçlamasında etkinliğinin belirlenmesi üzerinde bir araştırma

    Başlık çevirisi yok

    ALİ BAYAT

    Yüksek Lisans

    Türkçe

    Türkçe

    1987

    ZiraatÇukurova Üniversitesi

    Tarımsal Mekanizasyon Ana Bilim Dalı

    DOÇ. DR. YUSUF ZEREN

  4. Dört zamanlı türbaşarj direk püskürtmeli bir dizel motorunun bilgisayar ile sümülasyonu

    Computer aided simulation of a fourstroke turbocharged direct-anjection diesel engine

    MUSTAFA BALCI

    Doktora

    Türkçe

    Türkçe

    1986

    Makine MühendisliğiGazi Üniversitesi

    PROF. DR. OĞUZ BORAT