389 lines
32 KiB
TeX
389 lines
32 KiB
TeX
%! Author = alex
|
|
%! Date = 3/6/25
|
|
|
|
|
|
\section{Results}\label{sec:results}
|
|
%
|
|
%We begin by comparing the overall performance of all trained models and baselines using four key evaluation metrics:
|
|
%mean absolute error (MAE) and mean squared error (MSE), as well as the metrics for their respective sub-intervals.
|
|
%
|
|
%Overall, transformer-based models consistently outperformed LSTM variants and baseline methods across most evaluation criteria.
|
|
%Among the baselines, [e.g., "the rule-based method"] showed the weakest performance,
|
|
%while the [e.g., "windowed logistic regression"] performed competitively in certain contexts.
|
|
%Differences across models were most pronounced in MSE and \(R^2\),
|
|
%indicating that advanced architectures better captured higher-order dynamics and reduced large prediction errors.
|
|
%
|
|
%The previous section detailed the design and implementation of our fertility prediction pipeline,
|
|
%including data preprocessing, feature engineering, input encoding, and the development of several deep learning architectures.
|
|
%We now present the results of our evaluation, focusing on the predictive accuracy of the proposed models across different temporal resolutions,
|
|
%cycle types and use cases.
|
|
%Model performance is assessed using both overall metrics and biologically targeted subintervals,
|
|
%allowing for a nuanced comparison of approaches and their practical relevance to real-time fertility forecasting.
|
|
%Additionally, model performance is compared to the three baseline models introduced.
|
|
%
|
|
% provide information about the training behaviour and statistic of the different models??
|
|
|
|
\subsection{Overall Model Performance Across Architectures}\label{subsec:overall_model_performance_across_architectures}
|
|
|
|
|
|
% show why I selected the individual input configs for model config training
|
|
% selected by best mse fertility, use 2nd best, as it provides basically the same performance, but more input data for more complex model configs
|
|
|
|
\subsection{Fertility Probability Prediction}\label{subsec:fertility_probability_prediction}
|
|
|
|
\subsubsection{Impact of Input Window Length}\label{subsubsec:fert_impact_of_historical_context}
|
|
\begin{landscape}
|
|
\begin{table}
|
|
\small
|
|
\begin{tabularx}{\linewidth}{l*{6}{X}}
|
|
\toprule
|
|
\multirow{2}{*}{Input-Length in Days} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
|
|
\cmidrule(r){2-4} \cmidrule(r){5-7}
|
|
& Fertility Overall & Fertile Days & Non-Fertile Days & Fertility Overall & Fertile Days & Non-Fertile Days \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{LSTM}} \\ \midrule
|
|
10 & 0.0399 & 0.0857 & 0.0212 & 0.0045 & 0.0105 & 0.0021 \\
|
|
20 & 0.0421 & 0.0994 & 0.0180 & 0.0052 & 0.0144 & \underline{\textbf{0.0013}} \\
|
|
40 & 0.0406 & 0.0938 & 0.0184 & 0.0049 & 0.0128 & 0.0016 \\
|
|
80 & 0.0420 & 0.1006 & \underline{\textbf{0.0176}} & 0.0053 & 0.0147 & 0.0014 \\
|
|
160 & \underline{0.0378} & \underline{\textbf{0.0834}} & 0.0192 & \underline{0.0043} & \underline{\textbf{0.0102}} & 0.0019 \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{Transformer}} \\ \midrule
|
|
10 & 0.0443 & 0.0868 & 0.0273 & 0.0046 & 0.0110 & 0.0021 \\
|
|
20 & 0.0472 & 0.0949 & 0.0274 & 0.0051 & 0.0133 & 0.0018 \\
|
|
40 & 0.0413 & 0.0861 & 0.0233 & \underline{0.0044} & 0.0111 & 0.0017 \\
|
|
80 & 0.0437 & \underline{0.0859} & 0.0269 & 0.0045 & \underline{0.0108} & 0.0021 \\
|
|
160 & \underline{0.0411} & 0.0882 & \underline{0.0218} & 0.0045 & 0.0116 & \underline{0.0016} \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{Convolution-LSTM}} \\ \midrule
|
|
10 & 0.0435 & 0.0954 & 0.0224 & 0.0050 & 0.0133 & 0.0017 \\
|
|
20 & 0.0414 & 0.0996 & \underline{0.0179} & 0.0050 & 0.0145 & \underline{\textbf{0.0013}} \\
|
|
40 & \underline{0.0394} & \underline{0.0911} & 0.0184 & \underline{0.0045} & \underline{0.0122} & 0.0014 \\
|
|
80 & 0.0481 & 0.1005 & 0.0278 & 0.0054 & 0.0146 & 0.0019 \\
|
|
160 & 0.0505 & 0.1067 & 0.0290 & 0.0060 & 0.0166 & 0.0020 \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{Convolution-Transformer}} \\ \midrule
|
|
10 & 0.0459 & 0.0962 & 0.0256 & 0.0049 & 0.0136 & \underline{0.0015} \\
|
|
20 & 0.0423 & \underline{0.0848} & 0.0257 & 0.0045 & \underline{\textbf{0.0102}} & 0.0022 \\
|
|
40 & \underline{\textbf{0.0376}} & 0.0854 & \underline{0.0186} & \underline{\textbf{0.0041}} & 0.0108 & \underline{0.0015} \\
|
|
80 & 0.0414 & 0.0932 & 0.0210 & 0.0048 & 0.0128 & 0.0018 \\
|
|
160 & 0.0523 & 0.1018 & 0.0337 & 0.0059 & 0.0148 & 0.0026 \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Architectures and Input Lengths on a fixed Input Resolution of 12 Values per Day.
|
|
\underline{Underlined} values represent the best value for each metric within a model.
|
|
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
|
|
\label{tab:fertility_results_by_window_length}
|
|
\end{table}
|
|
\end{landscape}
|
|
|
|
Table~\ref{tab:fertility_results_by_window_length} summarises performance across metrics and input window lengths for the fertility-probability target.
|
|
|
|
For both the LSTM and Transformer, longer input windows generally yield better results.
|
|
The LSTM achieves the top scores in four of six metrics with the 160-day window, including the global best for MAE and MSE on both fertile and non-fertile days.
|
|
The only exception is non-fertile day errors, where shorter windows perform better.
|
|
|
|
The Transformer peaks at the 160-day window in three of six metrics but improves on non-fertile days with shorter inputs.
|
|
Across all window lengths, it is consistently outperformed by the LSTM\@.
|
|
|
|
Convolutional models perform best with medium-length windows (20--40 days).
|
|
The convolutional LSTM reaches its lowest errors for most metrics at 40 days, except for non-fertile MSE, where 20 days is optimal---also a global best.
|
|
The convolutional Transformer similarly favours 40 days for overall and non-fertile-day metrics, while fertile-day metrics perform best with 20-day inputs.
|
|
It achieves global best scores for 40-day MAE and MSE, and outperforms the convolutional LSTM in all but non-fertile day metrics.
|
|
|
|
\subsubsection{Impact of Input Resolution}\label{subsubsec:fert_impact_of_input_resolution}
|
|
Shifting focus from temporal span to sampling density, Table~\ref{tab:fertility_results_by_window_resolution} reports performance for the fertility-probability target across varying input resolutions.
|
|
Convolutional models are omitted, as their convolution layers inherently perform learnable resampling.
|
|
|
|
The LSTM outperforms the Transformer at all resolutions, with best-performing metrics scattered across the medium-to-high range (4--288 values/day) and no single optimum.
|
|
Fertile-day errors are lowest at 288 values/day, while non-fertile day errors peak at 12 values/day.
|
|
Overall MAE and MSE minima occur at 48 and 4 values/day, respectively.
|
|
|
|
The Transformer shows a clearer trend towards higher-resolution inputs.
|
|
Its fertile-day metrics are best at 4 values/day, whereas non-fertile day, overall MAE, and overall MSE scores peak between 48 and 288 values/day.
|
|
|
|
|
|
\begin{landscape}
|
|
\begin{table}
|
|
\small
|
|
\begin{tabularx}{\linewidth}{l*{6}{X}}
|
|
\toprule
|
|
\multirow{2}{*}{Values Per Day} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
|
|
\cmidrule(r){2-4} \cmidrule(r){5-7}
|
|
& Fertility Overall & Fertile Days & Non-Fertile Days & Fertility Overall & Fertile Days & Non-Fertile Days \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{LSTM}} \\
|
|
\midrule
|
|
1 & 0.0471 & 0.0987 & 0.0220 & 0.0062 & 0.0145 & 0.0022 \\
|
|
2 & 0.0457 & 0.0962 & 0.0231 & 0.0052 & 0.0133 & 0.0016 \\
|
|
4 & 0.0421 & 0.0898 & 0.0223 & \underline{\textbf{0.0046}}& 0.0116 & 0.0018 \\
|
|
12 & 0.0421 & 0.0994 & \underline{\textbf{0.0180}}& 0.0052 & 0.0144 & \underline{\textbf{0.0013}}\\
|
|
24 & 0.0419 & 0.0949 & 0.0201 & 0.0050 & 0.0132 & 0.0017 \\
|
|
48 & \underline{\textbf{0.0402}}& 0.0929 & 0.0182 & 0.0049 & 0.0126 & 0.0018 \\
|
|
72 & 0.0433 & 0.0972 & 0.0216 & 0.0052 & 0.0137 & 0.0018 \\
|
|
288 & 0.0410 & \underline{\textbf{0.0872}}& 0.0223 & 0.0049 & \underline{\textbf{0.0110}}& 0.0025 \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{Transformer}} \\
|
|
\midrule
|
|
1 & 0.0541 & 0.1015 & 0.0313 & 0.0063 & 0.0146 & 0.0023 \\
|
|
2 & 0.0468 & 0.0946 & 0.0259 & 0.0054 & 0.0128 & 0.0021 \\
|
|
4 & 0.0456 & \underline{0.0891}& 0.0277 & 0.0050 & \underline{0.0115}& 0.0023 \\
|
|
12 & 0.0472 & 0.0949 & 0.0274 & 0.0051 & 0.0133 & 0.0018 \\
|
|
24 & 0.0502 & 0.0960 & 0.0323 & 0.0052 & 0.0133 & 0.0021 \\
|
|
48 & 0.0447 & 0.0912 & \underline{0.0254}& \underline{0.0048}& 0.0122 & 0.0017 \\
|
|
72 & 0.0468 & 0.0975 & 0.0257 & 0.0051 & 0.0140 & \underline{0.0014} \\
|
|
288 & \underline{0.0449}& 0.0937 & 0.0257 & \underline{0.0048}& 0.0128 & 0.0017 \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Architectures and Input Resolutions on a fixed Input-Window-Length of 20 Days.
|
|
\underline{Underlined} values represent the best value for each metric within a model.
|
|
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
|
|
\label{tab:fertility_results_by_window_resolution}
|
|
\end{table}
|
|
\end{landscape}
|
|
|
|
\subsubsection{Impact of Model Parameters}\label{subsubsec:fert_impaoct_of_model_parameters}
|
|
\begin{landscape}
|
|
\begin{table}
|
|
\scriptsize
|
|
\begin{tabularx}{\linewidth}{l*{8}{X}}
|
|
\toprule
|
|
\multirow{2}{*}{Hidden Layer Size} &
|
|
\multirow{2}{*}{\# LSTM Layers} &
|
|
\multicolumn{3}{c}{MAE} &
|
|
\multicolumn{3}{c}{MSE} \\
|
|
\cmidrule(lr){3-5} \cmidrule(lr){6-8}
|
|
& & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
|
|
\midrule
|
|
16 & 1 & 0.0485 & 0.1050 & 0.0252 & 0.0059 & 0.0161 & 0.0016 \\
|
|
32 & 1 & 0.0477 & 0.1029 & 0.0248 & 0.0057 & 0.0155 & 0.0016 \\
|
|
32 & 2 & 0.0458 & 0.0922 & 0.0269 & 0.0052 & \textbf{0.0122} & 0.0024 \\
|
|
64 & 2 & 0.0447 & 0.0925 & 0.0250 & 0.0051 & 0.0123 & 0.0021 \\
|
|
128 & 2 & 0.0424 & 0.0951 & 0.0201 & 0.0051 & 0.0131 & 0.0018 \\
|
|
128 & 4 & 0.0427 & 0.0989 & 0.0191 & 0.0053 & 0.0145 & 0.0014 \\
|
|
256 & 4 & 0.0433 & 0.1050 & \textbf{0.0175} & 0.0055 & 0.0161 & \textbf{0.0011} \\
|
|
512 & 4 & \textbf{0.0399} & \textbf{0.0911} & 0.0189 & \textbf{0.0047} & \textbf{0.0122} & 0.0016 \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Parameters for the LSTM model with
|
|
a fixed input window length of 160 days and an input resolution of 12 values per day.
|
|
\textbf{Bold} values represent the best value for each metric within a model.}
|
|
\label{tab:fertility_results_by_model_parameters_lstm}
|
|
\end{table}
|
|
\begin{table}
|
|
\scriptsize
|
|
\begin{tabularx}{\linewidth}{l*{8}{X}}
|
|
\toprule
|
|
\multirow{2}{*}{Size of Embedding} &
|
|
\multirow{2}{*}{\# Encoder Layers} &
|
|
\multirow{2}{*}{\# Attention Heads} &
|
|
\multicolumn{3}{c}{MAE} &
|
|
\multicolumn{3}{c}{MSE} \\
|
|
\cmidrule(lr){4-6} \cmidrule(lr){7-9}
|
|
& & & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
|
|
\midrule
|
|
16 & 1 & 1 & 0.0708 & 0.1262 & 0.0475 & 0.0084 & 0.0227 & 0.0024 \\
|
|
32 & 1 & 1 & 0.0599 & 0.1163 & 0.0378 & 0.0070 & 0.0196 & 0.0019 \\
|
|
64 & 1 & 1 & 0.0601 & 0.1178 & 0.0373 & 0.0071 & 0.0200 & 0.0020 \\
|
|
64 & 2 & 2 & 0.0477 & 0.0908 & 0.0301 & 0.0049 & 0.0121 & 0.0020 \\
|
|
128 & 2 & 2 & 0.0449 & 0.0904 & 0.0263 & 0.0047 & 0.0119 & 0.0017 \\
|
|
128 & 4 & 4 & 0.0461 & 0.0979 & 0.0247 & 0.0050 & 0.0143 & 0.0013 \\
|
|
256 & 4 & 4 & 0.0475 & 0.0919 & 0.0293 & 0.0048 & 0.0125 & 0.0017 \\
|
|
512 & 4 & 4 & 0.0403 & \textbf{0.0831} & 0.0229 & \textbf{0.0043} & \textbf{0.0103} & 0.0018 \\
|
|
512 & 8 & 8 & \textbf{0.0395} & 0.0967 & \textbf{0.0159} & 0.0048 & 0.0139 & \textbf{0.0011} \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Parameters for the Transformer model with
|
|
a fixed input window length of 160 days and an input resolution of 12 values per day.
|
|
\textbf{Bold} values represent the best value for each metric within a model.}
|
|
\label{tab:fertility_results_by_model_parameters_transformer}
|
|
\end{table}
|
|
\end{landscape}
|
|
|
|
\subsubsection{Comparison with Baselines}\label{subsubsec:fert_comparison_with_baselines}
|
|
|
|
\subsection{Ovulation-Over Prediction}\label{subsubsec:ov_over_prediction}
|
|
|
|
\subsubsection{Impact of Input Window Length}\label{subsubsec:ov_over_impact_of_historical_context}
|
|
|
|
Table~\ref{tab:ov_over_results_by_window_length} presents performance across metrics and input window lengths for the ovulation-over target.
|
|
|
|
The LSTM and Transformer both outperform their convolutional counterparts.
|
|
The LSTM benefits from longer windows (160 days) except for the before-ovulation metric, where 20 days is optimal---also the \textbf{global best} across models.
|
|
The Transformer performs best with mid-length windows (40 days) for most metrics, with exceptions in after-ovulation performance (MAE: 80 days, MSE: 20 days).
|
|
|
|
Among convolutional models, the convolutional LSTM peaks at mid-length windows, reaching lowest MAE at 40 days and lowest MSE at 20 days.
|
|
The convolutional Transformer also favours 40 days overall, but before-ovulation performance benefits from longer inputs (160 days).
|
|
|
|
\begin{landscape}
|
|
\begin{table}
|
|
\small
|
|
\begin{tabularx}{\linewidth}{l*{6}{X}}
|
|
\toprule
|
|
\multirow{2}{*}{Input-Length in Days} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
|
|
\cmidrule(r){2-4} \cmidrule(r){5-7}
|
|
& OV-Over Overall & OV-Over Before OV & OV-Over After OV & OV-Over Overall & OV-Over Before OV & OV-Over After OV \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{LSTM}} \\
|
|
\midrule
|
|
10 & 0.1218 & 0.1044 & 0.1223 & 0.0612 & 0.0289 & 0.0711 \\
|
|
20 & 0.1153 & \underline{\textbf{0.0745}} & 0.1312 & 0.0641 & \underline{\textbf{0.0212}} & 0.0822 \\
|
|
40 & 0.1066 & 0.0820 & 0.1128 & 0.0616 & 0.0281 & 0.0740 \\
|
|
80 & 0.1173 & 0.0842 & 0.1291 & 0.0647 & 0.0263 & 0.0801 \\
|
|
160 & \underline{0.1039} & 0.1139 & \underline{0.0973} & \underline{0.0557} & 0.0462 & \underline{0.0580} \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{Transformer}} \\
|
|
\midrule
|
|
10 & 0.1137 & 0.1006 & 0.1186 & 0.0618 & 0.0336 & 0.0740 \\
|
|
20 & 0.1204 & 0.0771 & 0.1366 & 0.0690 & \underline{0.0255} & 0.0864 \\
|
|
40 & \underline{\textbf{0.1017}} & 0.1409 & \underline{\textbf{0.0883}} & \underline{\textbf{0.0533}} & 0.0621 & \underline{\textbf{0.0520}} \\
|
|
80 & 0.1138 & \underline{0.0897} & 0.1234 & 0.0654 & 0.0356 & 0.0788 \\
|
|
160 & 0.1076 & 0.0959 & 0.1141 & 0.0606 & 0.0379 & 0.0714 \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{Convolution-LSTM}} \\
|
|
\midrule
|
|
10 & 0.1847 & 0.1878 & 0.1840 & 0.0819 & 0.0581 & 0.0949 \\
|
|
20 & 0.1493 & 0.1274 & 0.1562 & \underline{0.0699} & \underline{0.0389} & \underline{0.0833} \\
|
|
40 & \underline{0.1455} & \underline{0.1089} & \underline{0.1561} & 0.0722 & 0.0358 & 0.0852 \\
|
|
80 & 0.2327 & 0.1796 & 0.2507 & 0.1168 & 0.0686 & 0.1327 \\
|
|
160 & 0.2518 & 0.2040 & 0.2757 & 0.1293 & 0.0878 & 0.1503 \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{Convolution-Transformer}} \\
|
|
\midrule
|
|
10 & 0.1472 & 0.1053 & 0.1651 & 0.0768 & 0.0317 & 0.0966 \\
|
|
20 & 0.1530 & 0.1307 & 0.1637 & 0.0745 & 0.0443 & 0.0886 \\
|
|
40 & \underline{0.1448} & 0.1228 & \underline{0.1514} & \underline{0.0709} & 0.0435 & \underline{0.0820} \\
|
|
80 & 0.1685 & 0.1164 & 0.1889 & 0.0865 & 0.0345 & 0.1089 \\
|
|
160 & 0.2440 & \underline{0.1051} & 0.3117 & 0.1403 & \underline{0.0286} & 0.1946 \\
|
|
\bottomrule
|
|
|
|
\end{tabularx}
|
|
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Architectures and Input Lengths on a fixed Input Resolution of 12 Values per Day.
|
|
\underline{Underlined} values represent the best value for each metric within a model.
|
|
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
|
|
\label{tab:ov_over_results_by_window_length}
|
|
\end{table}
|
|
\end{landscape}
|
|
|
|
\subsubsection{Impact of Input Resolution}\label{subsubsec:ov_over_impact_of_input_resolution}
|
|
|
|
Table~\ref{tab:ov_over_results_by_resolution} shows results for the ovulation-over target across different resolutions.
|
|
Convolutional models are excluded, as their convolution layers inherently perform resampling.
|
|
|
|
The LSTM generally performs best at medium resolutions (12--48 values/day), except for the after-ovulation metric, where higher resolutions yield better results.
|
|
The Transformer prefers higher resolutions, with optimal performance for all but the before-ovulation metric at 72 or 288 values/day.
|
|
Before-ovulation performance peaks at a medium resolution of 12 values/day.
|
|
|
|
\begin{landscape}
|
|
\begin{table}
|
|
\small
|
|
\begin{tabularx}{\linewidth}{l*{6}{X}}
|
|
\toprule
|
|
\multirow{2}{*}{Values Per Day} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
|
|
\cmidrule(r){2-4} \cmidrule(r){5-7}
|
|
& OV-Over Overall & OV-Over Before OV & OV-Over After OV & OV-Over Overall & OV-Over Before OV & OV-Over After OV \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{LSTM}} \\
|
|
\midrule
|
|
1 & 0.1691 & 0.1280 & 0.1819 & 0.0914 & 0.0411 & 0.1099 \\
|
|
2 & 0.1431 & 0.1239 & 0.1471 & 0.0747 & 0.0442 & 0.0849 \\
|
|
4 & 0.1344 & 0.1130 & 0.1417 & 0.0680 & 0.0371 & 0.0807 \\
|
|
12 & \underline{0.1153} & \underline{\textbf{0.0745}} & 0.1312 & 0.0641 & \underline{\textbf{0.0212}} & 0.0822 \\
|
|
24 & 0.1192 & 0.0975 & 0.1240 & \underline{0.0633} & 0.0288 & 0.0755 \\
|
|
48 & 0.1463 & 0.2493 & \underline{0.0973} & 0.0768 & 0.1203 & \underline{\textbf{0.0550}} \\
|
|
72 & 0.1587 & 0.1113 & 0.1844 & 0.0857 & 0.0341 & 0.1126 \\
|
|
288 & 0.1353 & 0.1801 & 0.1169 & 0.0726 & 0.0799 & 0.0713 \\
|
|
\midrule
|
|
\multicolumn{7}{c}{\textbf{Transformer}} \\
|
|
\midrule
|
|
1 & 0.1687 & 0.1582 & 0.1709 & 0.0883 & 0.0578 & 0.0996 \\
|
|
2 & 0.1469 & 0.0895 & 0.1676 & 0.0823 & 0.0260 & 0.1044 \\
|
|
4 & 0.1282 & 0.0999 & 0.1393 & 0.0704 & 0.0329 & 0.0862 \\
|
|
12 & 0.1204 & \underline{0.0771}& 0.1366 & 0.0690 & \underline{0.0255}& 0.0864 \\
|
|
24 & 0.1952 & 0.1649 & 0.2102 & 0.0904 & 0.0638 & 0.1040 \\
|
|
48 & 0.1137 & 0.1159 & 0.1073 & 0.0610 & 0.0480 & 0.0629 \\
|
|
72 & \underline{\textbf{0.1041}}& 0.1180 & \underline{\textbf{0.0947}}& \underline{\textbf{0.0585}}& 0.0538 & 0.0581 \\
|
|
288 & 0.1319 & 0.1879 & 0.1097 & 0.0617 & 0.0769 & \underline{0.0578}\\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Architectures and Input Resolutions on a fixed Input-Window-Length of 20 Days.
|
|
\underline{Underlined} values represent the best value for each metric within a model.
|
|
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
|
|
\label{tab:ov_over_results_by_resolution}
|
|
\end{table}
|
|
\end{landscape}
|
|
|
|
\subsubsection{Impact of Model Parameters}\label{subsubsec:ov_over_impaoct_of_model_parameters}
|
|
\begin{landscape}
|
|
\begin{table}
|
|
\scriptsize
|
|
\begin{tabularx}{\linewidth}{l*{8}{X}}
|
|
\toprule
|
|
\multirow{2}{*}{Hidden Layer Size} &
|
|
\multirow{2}{*}{\# LSTM Layers} &
|
|
\multicolumn{3}{c}{MAE} &
|
|
\multicolumn{3}{c}{MSE} \\
|
|
\cmidrule(lr){3-5} \cmidrule(lr){6-8}
|
|
& & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
|
|
\midrule
|
|
16 & 1 & 0.2253 & 0.1353 & 0.2735 & 0.1196 & 0.0397 & 0.1621 \\
|
|
32 & 1 & 0.1746 & 0.1255 & 0.2005 & 0.0860 & 0.0404 & 0.1109 \\
|
|
32 & 2 & 0.1349 & 0.1240 & 0.1399 & 0.0674 & 0.0396 & 0.0786 \\
|
|
64 & 2 & 0.1270 & 0.0943 & 0.1365 & 0.0645 & 0.0292 & 0.0772 \\
|
|
128 & 2 & 0.1151 & \textbf{0.0861} & 0.1243 & 0.0626 & \textbf{0.0274} & 0.0755 \\
|
|
128 & 4 & 0.1326 & 0.1064 & 0.1407 & 0.0658 & 0.0291 & 0.0811 \\
|
|
256 & 4 & 0.1353 & 0.1111 & 0.1407 & 0.0662 & 0.0360 & 0.0765 \\
|
|
512 & 4 & \textbf{0.1120} & 0.1358 & \textbf{0.0983} & \textbf{0.0616} & 0.0603 & \textbf{0.0613} \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Parameters for the LSTM model with
|
|
a fixed input window length of 160 days and an input resolution of 12 values per day.
|
|
\textbf{Bold} values represent the best value for each metric within a model.}
|
|
\label{tab:ov_over_results_by_model_parameters_lstm}
|
|
\end{table}
|
|
\begin{table}
|
|
\scriptsize
|
|
\begin{tabularx}{\linewidth}{l*{8}{X}}
|
|
\toprule
|
|
\multirow{2}{*}{Size of Embedding} &
|
|
\multirow{2}{*}{\# Encoder Layers} &
|
|
\multirow{2}{*}{\# Attention Heads} &
|
|
\multicolumn{3}{c}{MAE} &
|
|
\multicolumn{3}{c}{MSE} \\
|
|
\cmidrule(lr){4-6} \cmidrule(lr){7-9}
|
|
& & & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
|
|
\midrule
|
|
16 & 1 & 1 & 0.4792 & 0.5280 & 0.4675 & 0.2336 & 0.2820 & 0.2220 \\
|
|
32 & 1 & 1 & 0.4406 & 0.4572 & 0.4417 & 0.2046 & 0.2179 & 0.2068 \\
|
|
64 & 1 & 1 & 0.3614 & 0.3411 & 0.3796 & 0.1696 & 0.1604 & 0.1800 \\
|
|
64 & 2 & 2 & 0.1532 & \textbf{0.0810} & 0.1913 & 0.0904 & 0.0302 & 0.1214 \\
|
|
128 & 2 & 2 & 0.1476 & 0.0979 & 0.1739 & 0.0816 & 0.0332 & 0.1064 \\
|
|
128 & 4 & 4 & \textbf{0.1114} & 0.0887 & 0.1236 & 0.0668 & 0.0368 & 0.0821 \\
|
|
256 & 4 & 4 & 0.1310 & 0.0920 & 0.1433 & 0.0722 & \textbf{0.0293} & 0.0876 \\
|
|
512 & 4 & 4 & 0.1222 & 0.1047 & 0.1261 & 0.0610 & 0.0321 & 0.0708 \\
|
|
512 & 8 & 8 & 0.1126 & 0.1965 & \textbf{0.0753} & \textbf{0.0543} & 0.0868 & \textbf{0.0410} \\
|
|
\bottomrule
|
|
\end{tabularx}
|
|
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Parameters for the Transformer model with
|
|
a fixed input window length of 160 days and an input resolution of 12 values per day.
|
|
\textbf{Bold} values represent the best value for each metric within a model.}
|
|
\label{tab:ov_over_results_by_model_parameters_transformer}
|
|
\end{table}
|
|
\end{landscape}
|
|
|
|
\subsubsection{Comparison with Baselines}\label{subsubsec:ov_over_comparison_with_baselines}
|
|
|
|
\subsection{Stratified Analysis}\label{subsec:stratified_analysis}
|
|
|
|
\subsubsection{Regular vs Irregular Cycles}\label{subsubsec:regular_vs_irregular_cycles}
|
|
|
|
\subsubsection{Influence of User History Depth}\label{subsubsec:influence_of_past_user_data}
|
|
|
|
\subsection{Use-Case Evaluation Results}\label{subsec:use_case_evaluation_results}
|
|
|
|
\subsubsection{Contraception Use-Case Results}\label{subsubsec:use_case_contraception_results}
|
|
|
|
\subsubsection{Pregnancy Use-Case Results}\label{subsubsec:use_case_pregnancy_results}
|
|
|
|
\subsection{Summary of Key Findings}\label{subsec:summary_of_key_findings}
|