Files
temperature-based-fertility…/thesis/sections/results.tex
T
2025-08-06 15:14:19 +02:00

95 lines
4.7 KiB
TeX

%! Author = alex
%! Date = 3/6/25
\section{Results}\label{sec:results}
We begin by comparing the overall performance of all trained models and baselines using four key evaluation metrics:
mean absolute error (MAE), mean squared error (MSE), and the coefficient of determination (\(R^2\)) for the regression target,
as well as the metrics for their respective sub-intervals.
Figure~\ref{fig:results_performance_overview_by_model_type} provides an overview of these metrics across all model types.
Overall, transformer-based models consistently outperformed LSTM variants and baseline methods across most evaluation criteria.
Among the baselines, [e.g., "the rule-based method"] showed the weakest performance,
while the [e.g., "windowed logistic regression"] performed competitively in certain contexts.
Differences across models were most pronounced in MSE and \(R^2\),
indicating that advanced architectures better captured higher-order dynamics and reduced large prediction errors.
The previous section detailed the design and implementation of our fertility prediction pipeline,
including data preprocessing, feature engineering, input encoding, and the development of several deep learning architectures.
We now present the results of our evaluation, focusing on the predictive accuracy of the proposed models across different temporal resolutions,
cycle types and use cases.
Model performance is assessed using both overall metrics and biologically targeted subintervals,
allowing for a nuanced comparison of approaches and their practical relevance to real-time fertility forecasting.
Additionally, model performance is compared to the three baseline models introduced.
% provide information about the training behaviour and statistic of the different models??
\subsection{Overall Model Performance Across Architectures}\label{subsec:overall_model_performance_across_architectures}
\begin{figure}
\centering
\includegraphics[width=0.9\textwidth]{resources/figures/results/model_performance_overview}
\caption{Overview of the performances of all model types including the baselines on 4 selected performance metrics.}
\label{fig:results_performance_overview_by_model_type}
\end{figure}
\subsection{Fertility Probability Prediction Accuracy}\label{subsec:fertility_probability_precition_accuracy}
\begin{landscape}
\begin{table}[ht]
\centering
\caption{Model comparison for fertility prediction using MAE and MSE}
\begin{adjustbox}{max width=\linewidth}
\begin{tabular}{lllcccccc}
\toprule
\textbf{Model} & \textbf{Window} & \textbf{Daily} &
\textbf{MAE$_{fert}$} & \textbf{MAE$_{fert,during}$} & \textbf{MAE$_{fert,non}$} &
\textbf{MSE$_{fert}$} & \textbf{MSE$_{fert,during}$} & \textbf{MSE$_{fert,non}$} \\
\midrule
ModelA & 7 & Yes & 0.67 & 0.59 & 0.73 & 0.89 & 0.82 & 0.95 \\
ModelB & 14 & No & 0.65 & 0.58 & 0.71 & 0.87 & 0.80 & 0.93 \\
% More rows...
\bottomrule
\end{tabular}
\end{adjustbox}
\label{tab:fertility_comparison}
\end{table}
\end{landscape}
% show why I selected the individual input configs for model config training
% selected by best mse fertility, use 2nd best, as it provides basically the same performance, but more input data for more complex model configs
\subsubsection{Fertility Probability Prediction}\label{subsubsec:fertility_probability_prediction}
\subsubsection{Impact of Input Resolution}\label{subsubsec:fert_impact_of_input_resolution}
\subsubsection{Impact of Input Window Length}\label{subsubsec:fert_impact_of_historical_context}
\subsubsection{Comparison with Baselines}\label{subsubsec:fert_comparison_with_baselines}
\subsection{Ovulation-Over Prediction}\label{subsubsec:ov_over_prediction}
\subsubsection{Performance around Ovulation}\label{subsubsec:ov_over_performance_around_ovulation}
\subsubsection{Impact of Input Resolution}\label{subsubsec:ov_over_impact_of_input_resolution}
\subsubsection{Impact of Input Window Length}\label{subsubsec:ov_over_impact_of_historical_context}
\subsubsection{Comparison with Baselines}\label{subsubsec:ov_over_comparison_with_baselines}
\subsection{Stratified Analysis}\label{subsec:stratified_analysis}
\subsubsection{Regular vs Irregular Cycles}\label{subsubsec:regular_vs_irregular_cycles}
\subsubsection{Influence of User History Depth}\label{subsubsec:influence_of_past_user_data}
\subsection{Use-Case Evaluation Results}\label{subsec:use_case_evaluation_results}
\subsubsection{Contraception Use-Case Results}\label{subsubsec:use_case_contraception_results}
\subsubsection{Pregnancy Use-Case Results}\label{subsubsec:use_case_pregnancy_results}
\subsection{Summary of Key Findings}\label{subsec:summary_of_key_findings}