Files
temperature-based-fertility…/thesis/sections/results.tex
T
2025-08-18 16:48:27 +00:00

700 lines
52 KiB
TeX
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
%! Author = alex
%! Date = 3/6/25
\section{Results}\label{sec:results}
We summarize the main findings from our modeling experiments,
beginning with overall model performance across architectures,
followed by a detailed analysis of the individual architectures' performances for irregular and regular cycles and the use cases
introduced in the last section.
These results will also be compared with the three baseline methods to evaluate their performance versus less sophisticated methods.
We summarize the main findings from our modeling experiments, beginning with overall model performance across architectures.
We then provide a detailed analysis of model performance for irregular and regular cycles, as well as for
the pregnancy and contraception use-cases described earlier.
Additionally, these results will be compared to those of the three baseline methods to evaluate the benefit of more advanced modeling approaches.
\subsection{Overall Model Performance Across Architectures}\label{subsec:overall_model_performance_across_architectures}
We evaluated multiple model architectures to compare their effectiveness in predicting the fertility-probability and ovulation-over targets.
Based on these results, we selected the best-performing configuration for each architecture for further analysis.
\subsubsection{Fertility-Probability Prediction}\label{subsec:fertility_probability_prediction}
This section examines model performance in predicting the probability of fertility,
focusing on the effects of input window length, input resolution, and key architecture parameters.
\paragraph{Impact of Input Window Length:}\label{subsubsec:fert_impact_of_historical_context}
Table~\ref{tab:fertility_results_by_window_length} summarizes the effect of varying the input sequence length on model performance for the fertility-probability target.
Across all architectures, no single window length consistently outperformed others across all metrics.
For the \textbf{LSTM}, the longest input (160~days) yielded the lowest overall MAE (0.0378) and the best
fertile-day performance (MAE~$=0.0834$, MSE~$=0.0102$), while shorter sequences tended to perform slightly worse
, particularly for fertile-day prediction.
Non-fertile day performance was best at 20~days (MSE~$=0.0013$).
In the \textbf{Transformer}, the optimal MAE for overall fertility occurred at 160~days (0.0411),
but the lowest fertile-day error was achieved at 80~days (MAE~$=0.0859$).
The best non-fertile-day performance was seen with 160~days (MSE~$=0.0016$).
For the \textbf{Convolution-LSTM}, the shortest windows generally underperformed, with the best overall MAE (0.0394) and MSE (0.0045) obtained at 40~days.
Fertile-day metrics were optimal at 40~days as well, while non-fertile-day performance peaked at 20~days (MSE~$=0.0013$).
The \textbf{Convolution-Transformer} achieved the global best MAE (0.0376) and MSE (0.0041) for overall fertility at 40~days,
indicating that intermediate historical context was most effective for this architecture.
Fertile-day performance was strongest at 20~days (MSE~$=0.0102$), while non-fertile-day predictions benefited from shorter inputs (10 or 40~days).
Overall, results suggest that the optimal input length is architecture-dependent, with intermediate windows (20--40~days)
frequently yielding competitive or best performance, while extremely long sequences (160~days) only benefited certain architectures such as the LSTM\@.
\begin{landscape}
\begin{table}
\small
\begin{tabularx}{\linewidth}{l*{6}{X}}
\toprule
\multirow{2}{*}{Input-Length in Days} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
\cmidrule(r){2-4} \cmidrule(r){5-7}
& Fertility Overall & Fertile Days & Non-Fertile Days & Fertility Overall & Fertile Days & Non-Fertile Days \\
\midrule
\multicolumn{7}{c}{\textbf{LSTM}} \\ \midrule
10 & 0.0399 & 0.0857 & 0.0212 & 0.0045 & 0.0105 & 0.0021 \\
20 & 0.0421 & 0.0994 & 0.0180 & 0.0052 & 0.0144 & \underline{\textbf{0.0013}} \\
40 & 0.0406 & 0.0938 & 0.0184 & 0.0049 & 0.0128 & 0.0016 \\
80 & 0.0420 & 0.1006 & \underline{\textbf{0.0176}} & 0.0053 & 0.0147 & 0.0014 \\
160 & \underline{0.0378} & \underline{\textbf{0.0834}} & 0.0192 & \underline{0.0043} & \underline{\textbf{0.0102}} & 0.0019 \\
\midrule
\multicolumn{7}{c}{\textbf{Transformer}} \\ \midrule
10 & 0.0443 & 0.0868 & 0.0273 & 0.0046 & 0.0110 & 0.0021 \\
20 & 0.0472 & 0.0949 & 0.0274 & 0.0051 & 0.0133 & 0.0018 \\
40 & 0.0413 & 0.0861 & 0.0233 & \underline{0.0044} & 0.0111 & 0.0017 \\
80 & 0.0437 & \underline{0.0859} & 0.0269 & 0.0045 & \underline{0.0108} & 0.0021 \\
160 & \underline{0.0411} & 0.0882 & \underline{0.0218} & 0.0045 & 0.0116 & \underline{0.0016} \\
\midrule
\multicolumn{7}{c}{\textbf{Convolution-LSTM}} \\ \midrule
10 & 0.0435 & 0.0954 & 0.0224 & 0.0050 & 0.0133 & 0.0017 \\
20 & 0.0414 & 0.0996 & \underline{0.0179} & 0.0050 & 0.0145 & \underline{\textbf{0.0013}} \\
40 & \underline{0.0394} & \underline{0.0911} & 0.0184 & \underline{0.0045} & \underline{0.0122} & 0.0014 \\
80 & 0.0481 & 0.1005 & 0.0278 & 0.0054 & 0.0146 & 0.0019 \\
160 & 0.0505 & 0.1067 & 0.0290 & 0.0060 & 0.0166 & 0.0020 \\
\midrule
\multicolumn{7}{c}{\textbf{Convolution-Transformer}} \\ \midrule
10 & 0.0459 & 0.0962 & 0.0256 & 0.0049 & 0.0136 & \underline{0.0015} \\
20 & 0.0423 & \underline{0.0848} & 0.0257 & 0.0045 & \underline{\textbf{0.0102}} & 0.0022 \\
40 & \underline{\textbf{0.0376}} & 0.0854 & \underline{0.0186} & \underline{\textbf{0.0041}} & 0.0108 & \underline{0.0015} \\
80 & 0.0414 & 0.0932 & 0.0210 & 0.0048 & 0.0128 & 0.0018 \\
160 & 0.0523 & 0.1018 & 0.0337 & 0.0059 & 0.0148 & 0.0026 \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Architectures and Input Lengths on a fixed Input Resolution of 12 Values per Day.
\underline{Underlined} values represent the best value for each metric within a model.
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
\label{tab:fertility_results_by_window_length}
\end{table}
\end{landscape}
\paragraph{Impact of Input Resolution:}\label{subsubsec:fert_impact_of_input_resolution}
Table~\ref{tab:fertility_results_by_window_resolution} shows the effect of varying the input resolution (values per day)
on model performance for the fertility-probability target.
Note, that the convolutional models are not included here, as they have their own learned input representation via convolution.
No single resolution consistently outperformed others across all metrics, and optimal settings varied by architecture.
For the \textbf{LSTM}, the lowest overall MAE (0.0402) was obtained at 48~values/day, while the best overall MSE (0.0046) occurred at 4~values/day.
Fertile-day performance was optimal at 288~values/day (MAE~$=0.0872$, MSE~$=0.0110$), and non-fertile-day metrics were best at
12~values/day (MAE~$=0.0180$, MSE~$=0.0013$), both of which represent the global best values for these categories.
In the \textbf{Transformer}, the lowest overall MAE (0.0447) occurred at 48~values/day, while the best overall MSE (0.0048) was shared between 48~and 288~values/day.
Fertile-day performance peaked at 4~values/day (MAE~$=0.0891$, MSE~$=0.0115$),
whereas non-fertile-day metrics were best at 48~values/day (MAE~$=0.0254$) and 72~values/day (MSE~$=0.0014$).
Overall, the results indicate that intermediate input resolutions (4--48~values/day) often yielded the best overall performance,
while extreme resolutions (1 or 288~values/day) only benefited specific metrics such as fertile-day prediction for the LSTM\@.
\begin{landscape}
\begin{table}
\small
\begin{tabularx}{\linewidth}{l*{6}{X}}
\toprule
\multirow{2}{*}{Values Per Day} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
\cmidrule(r){2-4} \cmidrule(r){5-7}
& Fertility Overall & Fertile Days & Non-Fertile Days & Fertility Overall & Fertile Days & Non-Fertile Days \\
\midrule
\multicolumn{7}{c}{\textbf{LSTM}} \\
\midrule
1 & 0.0471 & 0.0987 & 0.0220 & 0.0062 & 0.0145 & 0.0022 \\
2 & 0.0457 & 0.0962 & 0.0231 & 0.0052 & 0.0133 & 0.0016 \\
4 & 0.0421 & 0.0898 & 0.0223 & \underline{\textbf{0.0046}}& 0.0116 & 0.0018 \\
12 & 0.0421 & 0.0994 & \underline{\textbf{0.0180}}& 0.0052 & 0.0144 & \underline{\textbf{0.0013}}\\
24 & 0.0419 & 0.0949 & 0.0201 & 0.0050 & 0.0132 & 0.0017 \\
48 & \underline{\textbf{0.0402}}& 0.0929 & 0.0182 & 0.0049 & 0.0126 & 0.0018 \\
72 & 0.0433 & 0.0972 & 0.0216 & 0.0052 & 0.0137 & 0.0018 \\
288 & 0.0410 & \underline{\textbf{0.0872}}& 0.0223 & 0.0049 & \underline{\textbf{0.0110}}& 0.0025 \\
\midrule
\multicolumn{7}{c}{\textbf{Transformer}} \\
\midrule
1 & 0.0541 & 0.1015 & 0.0313 & 0.0063 & 0.0146 & 0.0023 \\
2 & 0.0468 & 0.0946 & 0.0259 & 0.0054 & 0.0128 & 0.0021 \\
4 & 0.0456 & \underline{0.0891}& 0.0277 & 0.0050 & \underline{0.0115}& 0.0023 \\
12 & 0.0472 & 0.0949 & 0.0274 & 0.0051 & 0.0133 & 0.0018 \\
24 & 0.0502 & 0.0960 & 0.0323 & 0.0052 & 0.0133 & 0.0021 \\
48 & 0.0447 & 0.0912 & \underline{0.0254}& \underline{0.0048}& 0.0122 & 0.0017 \\
72 & 0.0468 & 0.0975 & 0.0257 & 0.0051 & 0.0140 & \underline{0.0014} \\
288 & \underline{0.0449}& 0.0937 & 0.0257 & \underline{0.0048}& 0.0128 & 0.0017 \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Architectures and Input Resolutions on a fixed Input-Window-Length of 20 Days.
\underline{Underlined} values represent the best value for each metric within a model.
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
\label{tab:fertility_results_by_window_resolution}
\end{table}
\end{landscape}
\paragraph{Impact of Model Parameters:}\label{subsubsec:fert_impaoct_of_model_parameters}
Tables~\ref{tab:fertility_results_by_model_parameters_lstm}--\ref{tab:fertility_results_by_model_parameters_conv_transformer} report
the results of the model parameter search across all architectures.
Each table shows the effect of varying hidden layer size and number of LSTM layers (for recurrent models),
or embedding size, number of encoder layers, and attention heads (for Transformer-based models).
For the \textbf{LSTM}, performance improved with increasing hidden layer size,
reaching the best overall values at 512 units with four layers (MAE~$=0.0399$, MSE~$=0.0047$). The lowest fertile-day errors were also
observed in this configuration, while non-fertile-day performance peaked at 256 units (MAE~$=0.0175$, MSE~$=0.0011$).
In the \textbf{Transformer}, larger embeddings and deeper networks generally improved performance.
The best overall MAE (0.0395) was achieved with a 512-dimensional embedding, eight encoder layers, and eight attention heads.
The lowest fertile-day errors occurred with a 512-dimensional embedding and four layers (MAE~$=0.0831$,
MSE~$=0.0103$), whereas non-fertile-day performance was strongest at 512~×~8 (MAE~$=0.0159$, MSE~$=0.0011$).
For the \textbf{Convolution-LSTM}, the best overall configuration was 256 hidden units with four layers,
yielding the lowest overall MSE (0.0042) and fertile-day MSE (0.0100).
Increasing to 512 units slightly reduced overall MAE (0.0380) and non-fertile-day MAE (0.0184).
In the \textbf{Convolution-Transformer}, the optimal configuration used a 512-dimensional embedding with four encoder layers and four attention heads,
achieving the best overall MAE (0.0377) and non-fertile-day performance (MAE~$=0.0177$, MSE~$=0.0013$).
Fertile-day prediction was strongest with eight encoder layers (MAE~$=0.0842$, MSE~$=0.0107$).
Taken together, these results indicate that larger model capacities generally improved performance across all architectures,
with the best configurations typically found at the higher end of the tested parameter ranges.
\begin{landscape}
\begin{table}
\scriptsize
\begin{tabularx}{\linewidth}{l*{8}{X}}
\midrule
\multicolumn{8}{c}{\textbf{LSTM}} \\
\toprule
\multirow{2}{*}{Hidden Layer Size} &
\multirow{2}{*}{\# LSTM Layers} &
\multicolumn{3}{c}{MAE} &
\multicolumn{3}{c}{MSE} \\
\cmidrule(lr){3-5} \cmidrule(lr){6-8}
& & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
\midrule
16 & 1 & 0.0485 & 0.1050 & 0.0252 & 0.0059 & 0.0161 & 0.0016 \\
32 & 1 & 0.0477 & 0.1029 & 0.0248 & 0.0057 & 0.0155 & 0.0016 \\
32 & 2 & 0.0458 & 0.0922 & 0.0269 & 0.0052 & \textbf{0.0122} & 0.0024 \\
64 & 2 & 0.0447 & 0.0925 & 0.0250 & 0.0051 & 0.0123 & 0.0021 \\
128 & 2 & 0.0424 & 0.0951 & 0.0201 & 0.0051 & 0.0131 & 0.0018 \\
128 & 4 & 0.0427 & 0.0989 & 0.0191 & 0.0053 & 0.0145 & 0.0014 \\
256 & 4 & 0.0433 & 0.1050 & \textbf{0.0175} & 0.0055 & 0.0161 & \textbf{0.0011} \\
512 & 4 & \textbf{0.0399} & \textbf{0.0911} & 0.0189 & \textbf{0.0047} & \textbf{0.0122} & 0.0016 \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Parameters for the LSTM model with
a fixed input window length of 160 days and an input resolution of 12 values per day.
\textbf{Bold} values represent the best value for each metric within a model.}
\label{tab:fertility_results_by_model_parameters_lstm}
\end{table}
\begin{table}
\scriptsize
\begin{tabularx}{\linewidth}{l*{8}{X}}
\midrule
\multicolumn{9}{c}{\textbf{Transformer}} \\
\toprule
\multirow{2}{*}{Size of Embedding} &
\multirow{2}{*}{\# Encoder Layers} &
\multirow{2}{*}{\# Attention Heads} &
\multicolumn{3}{c}{MAE} &
\multicolumn{3}{c}{MSE} \\
\cmidrule(lr){4-6} \cmidrule(lr){7-9}
& & & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
\midrule
16 & 1 & 1 & 0.0708 & 0.1262 & 0.0475 & 0.0084 & 0.0227 & 0.0024 \\
32 & 1 & 1 & 0.0599 & 0.1163 & 0.0378 & 0.0070 & 0.0196 & 0.0019 \\
64 & 1 & 1 & 0.0601 & 0.1178 & 0.0373 & 0.0071 & 0.0200 & 0.0020 \\
64 & 2 & 2 & 0.0477 & 0.0908 & 0.0301 & 0.0049 & 0.0121 & 0.0020 \\
128 & 2 & 2 & 0.0449 & 0.0904 & 0.0263 & 0.0047 & 0.0119 & 0.0017 \\
128 & 4 & 4 & 0.0461 & 0.0979 & 0.0247 & 0.0050 & 0.0143 & 0.0013 \\
256 & 4 & 4 & 0.0475 & 0.0919 & 0.0293 & 0.0048 & 0.0125 & 0.0017 \\
512 & 4 & 4 & 0.0403 & \textbf{0.0831} & 0.0229 & \textbf{0.0043} & \textbf{0.0103} & 0.0018 \\
512 & 8 & 8 & \textbf{0.0395} & 0.0967 & \textbf{0.0159} & 0.0048 & 0.0139 & \textbf{0.0011} \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Parameters for the Transformer model with
a fixed input window length of 160 days and an input resolution of 12 values per day.
\textbf{Bold} values represent the best value for each metric within a model.}
\label{tab:fertility_results_by_model_parameters_transformer}
\end{table}
\begin{table}
\scriptsize
\begin{tabularx}{\linewidth}{l*{8}{X}}
\midrule
\multicolumn{8}{c}{\textbf{Convolutional-LSTM}} \\
\toprule
\multirow{2}{*}{Hidden Layer Size} &
\multirow{2}{*}{\# LSTM Layers} &
\multicolumn{3}{c}{MAE} &
\multicolumn{3}{c}{MSE} \\
\cmidrule(lr){3-5} \cmidrule(lr){6-8}
& & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
\midrule
16 & 1 & 0.0487 & 0.1029 & 0.0264 & 0.0056 & 0.0154 & \textbf{0.0016} \\
32 & 1 & 0.0461 & 0.0956 & 0.0266 & 0.0050 & 0.0133 & 0.0017 \\
32 & 2 & 0.0448 & 0.0975 & 0.0232 & 0.0051 & 0.0139 & \textbf{0.0016} \\
64 & 2 & 0.0432 & 0.0921 & 0.0241 & 0.0048 & 0.0123 & 0.0019 \\
128 & 2 & 0.0400 & 0.0895 & 0.0204 & 0.0045 & 0.0119 & \textbf{0.0016} \\
128 & 4 & 0.0398 & 0.0871 & 0.0204 & 0.0044 & 0.0113 & \textbf{0.0016} \\
256 & 4 & 0.0389 & \textbf{0.0826} & 0.0209 & \textbf{0.0042} & \textbf{0.0100} & 0.0018 \\
512 & 4 & \textbf{0.0380} & 0.0869 & \textbf{0.0184} & 0.0044 & 0.0112 & 0.0017 \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Parameters for the convolutional LSTM model with
a fixed input window length of 40 days.
\textbf{Bold} values represent the best value for each metric within a model.}
\label{tab:fertility_results_by_model_parameters_conv_lstm}
\end{table}
\begin{table}
\scriptsize
\begin{tabularx}{\linewidth}{l*{8}{X}}
\midrule
\multicolumn{8}{c}{\textbf{Convolutional-Transformer}} \\
\toprule
\multirow{2}{*}{Size of Embedding} &
\multirow{2}{*}{\# Encoder Layers} &
\multirow{2}{*}{\# Attention Heads} &
\multicolumn{3}{c}{MAE} &
\multicolumn{3}{c}{MSE} \\
\cmidrule(lr){4-6} \cmidrule(lr){7-9}
& & & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
\midrule
16 & 1 & 1 & 0.0574 & 0.1005 & 0.0427 & 0.0061 & 0.0146 & 0.0032 \\
32 & 1 & 1 & 0.0462 & 0.0963 & 0.0266 & 0.0051 & 0.0135 & 0.0018 \\
64 & 1 & 1 & 0.0546 & 0.1006 & 0.0379 & 0.0059 & 0.0145 & 0.0027 \\
64 & 2 & 2 & 0.0427 & 0.0903 & 0.0237 & 0.0046 & 0.0118 & 0.0018 \\
128 & 2 & 2 & 0.0406 & 0.0868 & 0.0225 & 0.0044 & 0.0113 & 0.0017 \\
128 & 4 & 4 & 0.0411 & 0.0886 & 0.0230 & 0.0044 & 0.0116 & 0.0017 \\
256 & 4 & 4 & 0.0420 & 0.0866 & 0.0240 & \textbf{0.0043} & 0.0113 & 0.0016 \\
512 & 4 & 4 & \textbf{0.0377} & 0.0878 & \textbf{0.0177} & \textbf{0.0043} & 0.0117 & \textbf{0.0013} \\
512 & 8 & 8 & 0.0399 & \textbf{0.0842} & 0.0224 & 0.0044 & \textbf{0.0107} & 0.0019 \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Parameters for the convolutional Transformer model with
a fixed input window length of 40 days.
\textbf{Bold} values represent the best value for each metric within a model.}
\label{tab:fertility_results_by_model_parameters_conv_transformer}
\end{table}
\end{landscape}
\subsubsection{Ovulation-Over Prediction}\label{subsubsec:ov_over_prediction}
This section examines model performance in predicting the indicator for a passed ovulation, with attention to the influence of input window length,
the input resolution and the individual model architecture parameters.
\paragraph{Impact of Input Window Length:}\label{subsubsec:ov_over_impact_of_historical_context}
Table~\ref{tab:ov_over_results_by_window_length} reports the impact of input sequence length on performance for the ovulation-over target.
The optimal length varied across architectures and prediction phases.
For the \textbf{LSTM}, the best overall MAE (0.1039) and MSE (0.0557) were obtained with 160~days,
which also gave the lowest after-ovulation errors (MAE~$=0.0973$, MSE~$=0.0580$). However, the best before-ovulation
performance occurred at 20~days (MAE~$=0.0745$, MSE~$=0.0212$), which were the global bests for this phase.
In the \textbf{Transformer}, 40~days yielded the lowest overall MAE (0.1017) and MSE (0.0533),
as well as the lowest after-ovulation errors (MAE~$=0.0883$, MSE~$=0.0520$), all of which were global bests.
The best before-ovulation results were achieved at 80~days (MAE~$=0.0897$) and 20~days (MSE~$=0.0255$).
For the \textbf{Convolution-LSTM}, the shortest effective length was 40~days,
which achieved the lowest overall MAE (0.1455) and before-ovulation MAE (0.1089).
The best overall MSE (0.0699) and before-ovulation MSE (0.0389) were observed at 20~days.
After-ovulation errors were smallest at 20~days (MSE~$=0.0833$) and 40~days (MAE~$=0.1561$).
In the \textbf{Convolution-Transformer}, the best overall MAE (0.1448) and MSE (0.0709) occurred at 40~days,
which also minimized after-ovulation MAE (0.1514) and MSE (0.0820).
The best before-ovulation performance came from 160~days for MAE (0.1051) and 10~days for MSE (0.0317).
Overall, intermediate input lengths (20--40~days) were often optimal, particularly for before-ovulation prediction,
while longer sequences (160~days) occasionally improved after-ovulation accuracy.
\begin{landscape}
\begin{table}
\small
\begin{tabularx}{\linewidth}{l*{6}{X}}
\toprule
\multirow{2}{*}{Input-Length in Days} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
\cmidrule(r){2-4} \cmidrule(r){5-7}
& OV-Over Overall & OV-Over Before OV & OV-Over After OV & OV-Over Overall & OV-Over Before OV & OV-Over After OV \\
\midrule
\multicolumn{7}{c}{\textbf{LSTM}} \\
\midrule
10 & 0.1218 & 0.1044 & 0.1223 & 0.0612 & 0.0289 & 0.0711 \\
20 & 0.1153 & \underline{\textbf{0.0745}} & 0.1312 & 0.0641 & \underline{\textbf{0.0212}} & 0.0822 \\
40 & 0.1066 & 0.0820 & 0.1128 & 0.0616 & 0.0281 & 0.0740 \\
80 & 0.1173 & 0.0842 & 0.1291 & 0.0647 & 0.0263 & 0.0801 \\
160 & \underline{0.1039} & 0.1139 & \underline{0.0973} & \underline{0.0557} & 0.0462 & \underline{0.0580} \\
\midrule
\multicolumn{7}{c}{\textbf{Transformer}} \\
\midrule
10 & 0.1137 & 0.1006 & 0.1186 & 0.0618 & 0.0336 & 0.0740 \\
20 & 0.1204 & 0.0771 & 0.1366 & 0.0690 & \underline{0.0255} & 0.0864 \\
40 & \underline{\textbf{0.1017}} & 0.1409 & \underline{\textbf{0.0883}} & \underline{\textbf{0.0533}} & 0.0621 & \underline{\textbf{0.0520}} \\
80 & 0.1138 & \underline{0.0897} & 0.1234 & 0.0654 & 0.0356 & 0.0788 \\
160 & 0.1076 & 0.0959 & 0.1141 & 0.0606 & 0.0379 & 0.0714 \\
\midrule
\multicolumn{7}{c}{\textbf{Convolution-LSTM}} \\
\midrule
10 & 0.1847 & 0.1878 & 0.1840 & 0.0819 & 0.0581 & 0.0949 \\
20 & 0.1493 & 0.1274 & 0.1562 & \underline{0.0699} & \underline{0.0389} & \underline{0.0833} \\
40 & \underline{0.1455} & \underline{0.1089} & \underline{0.1561} & 0.0722 & 0.0358 & 0.0852 \\
80 & 0.2327 & 0.1796 & 0.2507 & 0.1168 & 0.0686 & 0.1327 \\
160 & 0.2518 & 0.2040 & 0.2757 & 0.1293 & 0.0878 & 0.1503 \\
\midrule
\multicolumn{7}{c}{\textbf{Convolution-Transformer}} \\
\midrule
10 & 0.1472 & 0.1053 & 0.1651 & 0.0768 & 0.0317 & 0.0966 \\
20 & 0.1530 & 0.1307 & 0.1637 & 0.0745 & 0.0443 & 0.0886 \\
40 & \underline{0.1448} & 0.1228 & \underline{0.1514} & \underline{0.0709} & 0.0435 & \underline{0.0820} \\
80 & 0.1685 & 0.1164 & 0.1889 & 0.0865 & 0.0345 & 0.1089 \\
160 & 0.2440 & \underline{0.1051} & 0.3117 & 0.1403 & \underline{0.0286} & 0.1946 \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Architectures and Input Lengths on a fixed Input Resolution of 12 Values per Day.
\underline{Underlined} values represent the best value for each metric within a model.
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
\label{tab:ov_over_results_by_window_length}
\end{table}
\end{landscape}
\paragraph{Impact of Input Resolution:}\label{subsubsec:ov_over_impact_of_input_resolution}
Table~\ref{tab:ov_over_results_by_resolution} presents the effect of varying input resolution (values per day)
on performance for the ovulation-over target.
Note, that the convolutional models are not included here, as they have their own learned input representation via convolution.
The best-performing resolution differed across architectures and prediction phases.
For the \textbf{LSTM}, the lowest overall MAE (0.1153) was achieved at 12~values/day,
which also produced the global best before-ovulation results (MAE~$=0.0745$, MSE~$=0.0212$).
The best overall MSE (0.0633) was observed at 24~values/day.
After-ovulation performance was strongest at 48~values/day (MAE~$=0.0973$, MSE~$=0.0550$), the latter being a global best.
In the \textbf{Transformer}, the optimal overall MAE (0.1041) and MSE (0.0585) were both achieved at 72~values/day, which also yielded the
lowest after-ovulation MAE (0.0947), all of which were global bests.
The best before-ovulation MAE (0.0771) and MSE (0.0255) were found at 12~values/day.
The lowest after-ovulation MSE (0.0578) occurred at 288~values/day.
Overall, intermediate input resolutions (12--72~values/day) tended to perform best for ovulation-over prediction,
with 12~values/day favouring before-ovulation performance and 48--72~values/day improving after-ovulation accuracy.
\begin{landscape}
\begin{table}
\small
\begin{tabularx}{\linewidth}{l*{6}{X}}
\toprule
\multirow{2}{*}{Values Per Day} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
\cmidrule(r){2-4} \cmidrule(r){5-7}
& OV-Over Overall & OV-Over Before OV & OV-Over After OV & OV-Over Overall & OV-Over Before OV & OV-Over After OV \\
\midrule
\multicolumn{7}{c}{\textbf{LSTM}} \\
\midrule
1 & 0.1691 & 0.1280 & 0.1819 & 0.0914 & 0.0411 & 0.1099 \\
2 & 0.1431 & 0.1239 & 0.1471 & 0.0747 & 0.0442 & 0.0849 \\
4 & 0.1344 & 0.1130 & 0.1417 & 0.0680 & 0.0371 & 0.0807 \\
12 & \underline{0.1153} & \underline{\textbf{0.0745}} & 0.1312 & 0.0641 & \underline{\textbf{0.0212}} & 0.0822 \\
24 & 0.1192 & 0.0975 & 0.1240 & \underline{0.0633} & 0.0288 & 0.0755 \\
48 & 0.1463 & 0.2493 & \underline{0.0973} & 0.0768 & 0.1203 & \underline{\textbf{0.0550}} \\
72 & 0.1587 & 0.1113 & 0.1844 & 0.0857 & 0.0341 & 0.1126 \\
288 & 0.1353 & 0.1801 & 0.1169 & 0.0726 & 0.0799 & 0.0713 \\
\midrule
\multicolumn{7}{c}{\textbf{Transformer}} \\
\midrule
1 & 0.1687 & 0.1582 & 0.1709 & 0.0883 & 0.0578 & 0.0996 \\
2 & 0.1469 & 0.0895 & 0.1676 & 0.0823 & 0.0260 & 0.1044 \\
4 & 0.1282 & 0.0999 & 0.1393 & 0.0704 & 0.0329 & 0.0862 \\
12 & 0.1204 & \underline{0.0771}& 0.1366 & 0.0690 & \underline{0.0255}& 0.0864 \\
24 & 0.1952 & 0.1649 & 0.2102 & 0.0904 & 0.0638 & 0.1040 \\
48 & 0.1137 & 0.1159 & 0.1073 & 0.0610 & 0.0480 & 0.0629 \\
72 & \underline{\textbf{0.1041}}& 0.1180 & \underline{\textbf{0.0947}}& \underline{\textbf{0.0585}}& 0.0538 & 0.0581 \\
288 & 0.1319 & 0.1879 & 0.1097 & 0.0617 & 0.0769 & \underline{0.0578}\\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Architectures and Input Resolutions on a fixed Input-Window-Length of 20 Days.
\underline{Underlined} values represent the best value for each metric within a model.
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
\label{tab:ov_over_results_by_resolution}
\end{table}
\end{landscape}
\paragraph{Impact of Model Parameters:}\label{subsubsec:ov_over_impaoct_of_model_parameters}
Tables~\ref{tab:ov_over_results_by_model_parameters_lstm}--\ref{tab:ov_over_results_by_model_parameters_conv_transformer}
show the results of the parameter exploration for the ovulation-over target.
Each table reports the effect of varying hidden layer size and number of LSTM layers (for recurrent models),
or embedding size, number of encoder layers, and number of attention heads (for Transformer-based models).
For the \textbf{LSTM}, performance improved with increasing hidden layer size,
with the best overall MAE (0.1120) and MSE (0.0616) obtained at 512 units with four layers.
This configuration also yielded the lowest after-ovulation errors (MAE~$=0.0983$, MSE~$=0.0613$).
The lowest before- ovulation errors were observed at 128 units with two layers (MAE~$=0.0861$, MSE~$=0.0274$).
In the \textbf{Transformer}, smaller configurations performed poorly, while larger ones markedly improved results.
The best overall MAE (0.1114) was achieved with a 128-dimensional embedding and four encoder layers,
whereas the best overall MSE (0.0543) occurred with a 512-dimensional embedding and eight encoder layers.
The lowest before-ovulation errors were found at 64 dimensions with two layers (MAE~$=0.0810$, MSE~$=0.0302$),
while after-ovulation performance was best at 512 dimensions with eight layers (MAE~$=0.0753$, MSE~$=0.0410$).
For the \textbf{Convolution-LSTM}, the best overall configuration used 256 hidden units with four layers,
reaching the lowest overall MAE (0.1424) and MSE (0.0687).
This configuration also minimized after-ovulation errors (MAE~$=0.1369$, MSE~$=0.0715$).
Before-ovulation performance was strongest with 128 units and two layers (MAE~$=0.1138$, MSE~$=0.0357$).
In the \textbf{Convolution-Transformer}, the best overall MAE (0.1495) and MSE (0.0703)
were achieved with a 512-dimensional embedding and four encoder layers.
This configuration also gave the lowest after-ovulation errors (MAE~$=0.1551$, MSE~$=0.0814$).
Before-ovulation errors were lowest at 512 dimensions with eight layers (MAE~$=0.1025$, MSE~$=0.0339$).
Overall, the results show that larger configurations generally improved performance across architectures for the ovulation-over target,
with the best outcomes typically found at the higher-capacity settings.
\begin{landscape}
\begin{table}
\scriptsize
\begin{tabularx}{\linewidth}{l*{8}{X}}
\midrule
\multicolumn{8}{c}{\textbf{LSTM}} \\
\toprule
\multirow{2}{*}{Hidden Layer Size} &
\multirow{2}{*}{\# LSTM Layers} &
\multicolumn{3}{c}{MAE} &
\multicolumn{3}{c}{MSE} \\
\cmidrule(lr){3-5} \cmidrule(lr){6-8}
& & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
\midrule
16 & 1 & 0.2253 & 0.1353 & 0.2735 & 0.1196 & 0.0397 & 0.1621 \\
32 & 1 & 0.1746 & 0.1255 & 0.2005 & 0.0860 & 0.0404 & 0.1109 \\
32 & 2 & 0.1349 & 0.1240 & 0.1399 & 0.0674 & 0.0396 & 0.0786 \\
64 & 2 & 0.1270 & 0.0943 & 0.1365 & 0.0645 & 0.0292 & 0.0772 \\
128 & 2 & 0.1151 & \textbf{0.0861} & 0.1243 & 0.0626 & \textbf{0.0274} & 0.0755 \\
128 & 4 & 0.1326 & 0.1064 & 0.1407 & 0.0658 & 0.0291 & 0.0811 \\
256 & 4 & 0.1353 & 0.1111 & 0.1407 & 0.0662 & 0.0360 & 0.0765 \\
512 & 4 & \textbf{0.1120} & 0.1358 & \textbf{0.0983} & \textbf{0.0616} & 0.0603 & \textbf{0.0613} \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Parameters for the LSTM model with
a fixed input window length of 160 days and an input resolution of 12 values per day.
\textbf{Bold} values represent the best value for each metric within a model.}
\label{tab:ov_over_results_by_model_parameters_lstm}
\end{table}
\begin{table}
\scriptsize
\begin{tabularx}{\linewidth}{l*{8}{X}}
\midrule
\multicolumn{9}{c}{\textbf{Transformer}} \\
\toprule
\multirow{2}{*}{Size of Embedding} &
\multirow{2}{*}{\# Encoder Layers} &
\multirow{2}{*}{\# Attention Heads} &
\multicolumn{3}{c}{MAE} &
\multicolumn{3}{c}{MSE} \\
\cmidrule(lr){4-6} \cmidrule(lr){7-9}
& & & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
\midrule
16 & 1 & 1 & 0.4792 & 0.5280 & 0.4675 & 0.2336 & 0.2820 & 0.2220 \\
32 & 1 & 1 & 0.4406 & 0.4572 & 0.4417 & 0.2046 & 0.2179 & 0.2068 \\
64 & 1 & 1 & 0.3614 & 0.3411 & 0.3796 & 0.1696 & 0.1604 & 0.1800 \\
64 & 2 & 2 & 0.1532 & \textbf{0.0810} & 0.1913 & 0.0904 & 0.0302 & 0.1214 \\
128 & 2 & 2 & 0.1476 & 0.0979 & 0.1739 & 0.0816 & 0.0332 & 0.1064 \\
128 & 4 & 4 & \textbf{0.1114} & 0.0887 & 0.1236 & 0.0668 & 0.0368 & 0.0821 \\
256 & 4 & 4 & 0.1310 & 0.0920 & 0.1433 & 0.0722 & \textbf{0.0293} & 0.0876 \\
512 & 4 & 4 & 0.1222 & 0.1047 & 0.1261 & 0.0610 & 0.0321 & 0.0708 \\
512 & 8 & 8 & 0.1126 & 0.1965 & \textbf{0.0753} & \textbf{0.0543} & 0.0868 & \textbf{0.0410} \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Parameters for the Transformer model with
a fixed input window length of 160 days and an input resolution of 12 values per day.
\textbf{Bold} values represent the best value for each metric within a model.}
\label{tab:ov_over_results_by_model_parameters_transformer}
\end{table}
\begin{table}
\scriptsize
\begin{tabularx}{\linewidth}{l*{8}{X}}
\midrule
\multicolumn{8}{c}{\textbf{Convolutional-LSTM}} \\
\toprule
\multirow{2}{*}{Hidden Layer Size} &
\multirow{2}{*}{\# LSTM Layers} &
\multicolumn{3}{c}{MAE} &
\multicolumn{3}{c}{MSE} \\
\cmidrule(lr){3-5} \cmidrule(lr){6-8}
& & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
\midrule
16 & 1 & 0.2920 & 0.2198 & 0.3344 & 0.1264 & 0.0602 & 0.1653 \\
32 & 1 & 0.2141 & 0.1598 & 0.2377 & 0.0912 & 0.0398 & 0.1149 \\
32 & 2 & 0.1990 & 0.1556 & 0.2135 & 0.0852 & 0.0419 & 0.1025 \\
64 & 2 & 0.1743 & 0.1345 & 0.1845 & 0.0789 & 0.0400 & 0.0923 \\
128 & 2 & 0.1524 & 0.1138 & 0.1633 & 0.0717 & \textbf{0.0357} & 0.0845 \\
128 & 4 & 0.1579 & 0.1185 & 0.1670 & 0.0777 & 0.0401 & 0.0894 \\
256 & 4 & \textbf{0.1424} & 0.1425 & \textbf{0.1369} & \textbf{0.0687} & 0.0546 & \textbf{0.0715} \\
512 & 4 & 0.1436 & \textbf{0.1166} & 0.1523 & 0.0699 & 0.0382 & 0.0820 \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Parameters for the convolutional LSTM model with
a fixed input window length of 40 days.
\textbf{Bold} values represent the best value for each metric within a model.}
\label{tab:ov_over_results_by_model_parameters_conv_lstm}
\end{table}
\begin{table}
\scriptsize
\begin{tabularx}{\linewidth}{l*{8}{X}}
\midrule
\multicolumn{8}{c}{\textbf{Convolutional-Transformer}} \\
\toprule
\multirow{2}{*}{Size of Embedding} &
\multirow{2}{*}{\# Encoder Layers} &
\multirow{2}{*}{\# Attention Heads} &
\multicolumn{3}{c}{MAE} &
\multicolumn{3}{c}{MSE} \\
\cmidrule(lr){4-6} \cmidrule(lr){7-9}
& & & Fert Overall & Fert Days & Non-Fert Days & Fert Overall & Fert Days & Non-Fert Days \\
\midrule
16 & 1 & 1 & 0.3759 & 0.3602 & 0.3855 & 0.1597 & 0.1498 & 0.1659 \\
32 & 1 & 1 & 0.1757 & 0.1119 & 0.2017 & 0.0916 & 0.0314 & 0.1167 \\
64 & 1 & 1 & 0.2771 & 0.2395 & 0.2978 & 0.1303 & 0.1133 & 0.1402 \\
64 & 2 & 2 & 0.1610 & 0.1149 & 0.1777 & 0.0816 & \textbf{0.0310} & 0.1021 \\
128 & 2 & 2 & 0.1501 & 0.0997 & 0.1725 & 0.0783 & 0.0316 & 0.0991 \\
128 & 4 & 4 & 0.1668 & 0.1472 & 0.1758 & 0.0801 & 0.0488 & 0.0953 \\
256 & 4 & 4 & 0.1529 & 0.1317 & 0.1597 & 0.0742 & 0.0435 & 0.0861 \\
512 & 4 & 4 & \textbf{0.1495} & 0.1298 & \textbf{0.1551} & \textbf{0.0703} & 0.0409 & \textbf{0.0814} \\
512 & 8 & 8 & 0.1576 & \textbf{0.1025} & 0.1848 & 0.0825 & 0.0339 & 0.1066 \\
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Parameters for the convolutional Transformer model with
a fixed input window length of 40 days.
\textbf{Bold} values represent the best value for each metric within a model.}
\label{tab:ov_over_results_by_model_parameters_conv_transformer}
\end{table}
\end{landscape}
\subsubsection{Best Model Configuration Selection}\label{subsubsec:results_best_model_config_selection}
Following the selection procedure described in Section~\ref{subsubsec:methodology_best_model_config_selection}
the best configuration for each model architecture was identified based on the Fertility-Overall MSE and, where applicable,
the general tendencies of the model.
Table~\ref{tab:best_configs_lstm} and~\ref{tab:best_configs_transformer} summarize the selected input window length,
input resolution, and model complexity for each architecture.
These configurations are used in all subsequent experiments, including the irregular cycles analysis and the use case evaluation.
\begin{table}[htbp]
\centering
\scriptsize
\begin{tabularx}{\linewidth}{lXXXX}
\toprule
\textbf{Architecture} &
\textbf{Input Window Length} &
\textbf{Input Resolution} &
\textbf{Hidden Layer Size} &
\textbf{\# LSTM Layers} \\
\midrule
LSTM & 160 & 12 & 512 & 4 \\
Convolution-LSTM & 40 & 288 & 512 & 4 \\
\bottomrule
\end{tabularx}
\caption{Selected configurations for LSTM-based architectures. Input Window Length is given in days and input resolution in values per day.}
\label{tab:best_configs_lstm}
\end{table}
\begin{table}[htbp]
\centering
\scriptsize
\begin{tabularx}{\linewidth}{lXXXXX}
\toprule
\textbf{Architecture} &
\textbf{Input Window Length} &
\textbf{Input Resolution} &
\textbf{Embedding Size} &
\textbf{\# Encoder Layers} &
\textbf{\# Attention Heads} \\
\midrule
Transformer & 160 & 12 & 512 & 4 & 4 \\
Convolution-Transformer & 40 & 288 & 512 & 4 & 4 \\
\bottomrule
\end{tabularx}
\caption{Selected configurations for Transformer-based architectures. Input Window Length is given in days and input resolution in values per day.}
\label{tab:best_configs_transformer}
\end{table}
\subsection{Stratified Analysis}\label{subsec:stratified_analysis}
\subsubsection{Influence of User History Depth}\label{subsubsec:influence_of_past_user_data}
% don't forget to also add baseline to tables
\subsubsection{Regular vs Irregular Cycles}\label{subsubsec:regular_vs_irregular_cycles}
\begin{landscape}
\begin{table}
\small
\begin{tabularx}{\linewidth}{l*{6}{X}}
\toprule
\multirow{2}{*}{Model Architecture} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
\cmidrule(r){2-4} \cmidrule(r){5-7}
& Fertility Overall & Fertile Days & Non-Fertile Days & Fertility Overall & Fertile Days & Non-Fertile Days \\
\midrule
\multicolumn{7}{c}{\textbf{Regular Cycle Group}} \\
\midrule
\midrule
\multicolumn{7}{c}{\textbf{Irregular Cycle Group}} \\
\midrule
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Fertility-Probability Target across Different Model Architectures for the Regular and Irregular Cycle Groups.
\underline{Underlined} values represent the best value for each metric within a model.
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
\label{tab:regular_vs_irregular_fertility_results}
\end{table}
\end{landscape}
\begin{landscape}
\begin{table}
\small
\begin{tabularx}{\linewidth}{l*{6}{X}}
\toprule
\multirow{2}{*}{Model} & \multicolumn{3}{c}{MAE} & \multicolumn{3}{c}{MSE} \\
\cmidrule(r){2-4} \cmidrule(r){5-7}
& OV-Over Overall & OV-Over Before OV & OV-Over After OV & OV-Over Overall & OV-Over Before OV & OV-Over After OV \\
\midrule
\multicolumn{7}{c}{\textbf{Regular Cycle Group}} \\
\midrule
\midrule
\multicolumn{7}{c}{\textbf{Irregular Cycle Group}} \\
\midrule
\bottomrule
\end{tabularx}
\caption{Evaluation Metrics for the Ovulation-Over Target across Different Model Architectures for the Regular and Irregular Cycle Groups.
\underline{Underlined} values represent the best value for each metric within a model.
\textbf{\underline{Bold + Underlined}} values represent the global best values across all models for a given metric.}
\label{tab:regular_vs_irregular_ov_over_results}
\end{table}
\end{landscape}
\subsection{Use-Case Evaluation Results}\label{subsec:use_case_evaluation_results}
\subsubsection{Contraception Use-Case Results}\label{subsubsec:use_case_contraception_results}
\subsubsection{Pregnancy Use-Case Results}\label{subsubsec:use_case_pregnancy_results}
\subsection{Summary of Key Findings}\label{subsec:summary_of_key_findings}