Skip to content
AI.info

Research

Augmented Web Usage Mining and User Experience Optimization with CAWAL's Enriched Analytics Data

Overview Research area: Human-Computer Interaction, specifically Web Usage Mining (WUM) and user experience (UX) optimization for large enterprise web portals. Technical level: Intermediate. The paper

Augmented Web Usage Mining and User Experience Optimization with CAWAL's Enriched Analytics Data
arXiv
2510.17253
Published
2025-10-20
Authors
Özkan Canay, {Ü}mit Kocabıcak

AI summary

Overview

Research area: Human-Computer Interaction, specifically Web Usage Mining (WUM) and user experience (UX) optimization for large enterprise web portals.

Technical level: Intermediate. The paper assumes familiarity with web analytics concepts (sessions, pageviews, bounce rate, logs) and with data mining techniques such as association rule mining and chi-squared testing, but it explains the CAWAL framework and the AWUM process from first principles.

Scope: The paper introduces Augmented Web Usage Mining (AWUM), a methodology that uses enriched analytical data produced by the CAWAL framework — rather than raw web server logs — to analyze user behavior and support UX optimization on a large multi-service university portal.

What This Paper Is About

Traditional web usage mining depends on raw web server logs, which are semi-structured, incomplete about in-page and cross-service interactions, and require a costly pre-processing phase of cleaning, transformation, and user/session identification. This paper proposes pulling already-structured, enriched session and pageview data from the CAWAL (Combined Application Log and Web Analytics) framework and feeding it directly into mining, skipping pre-processing entirely. The goal is to show that this "augmented" pipeline yields more accurate and efficient analysis of user behavior — and therefore better UX decisions — on a large-scale, load-balanced, multi-service web portal.

Key Contributions

  1. Introduces the AWUM (Augmented Web Usage Mining) approach, which integrates enriched analytics data from the CAWAL framework to analyze user behavior more accurately and in more detail than traditional methods.
  2. Shows that CAWAL simplifies the WUM process by removing the need for extensive pre-processing, delivering structured, high-quality data directly from application logs and web analytics.
  3. Demonstrates that CAWAL's enriched datasets can improve the accuracy and efficiency of machine learning, user behavior prediction, and classification models.
  4. Contributes UX optimization insights through deep analysis of user interactions, intended to guide web portal performance improvements and strategic decisions across web services.

Main Findings

  • Volume of data processed: More than 1.2 million session records (specifically 1,220,916 records) collected over one month by CAWAL were processed and transformed into 8.5 GB of enriched data.
  • Multi-page dominance: Of 1,220,916 total sessions, 156,707 were single-page and 1,064,209 were multi-page. Single-page sessions constitute 12.84% of sessions and multi-page sessions 87.16%.
  • Pageview concentration: Single-page sessions produced 156,707 pageviews (1.95% of total pageviews) while multi-page sessions produced 7,882,632 pageviews (98.05% of total pageviews).
  • Browser and referrer differences are statistically significant: The chi-squared test gave Browser_Type χ² = 68.19, p = 0.00, DoF = 2, and Referer_Type χ² = 38.72, p = 0.00, DoF = 5 — indicating significant relationships with session type.
  • Language and location differences are not significant: User_Language_TR gave χ² = 0.00, p = 1.00, DoF = 1, and User_Location gave χ² = 0.12, p = 0.94, DoF = 2, indicating no significant relationship with session type.
  • Standard browsers dominate multi-page sessions: For Browser_Type, 1-Standard Browser accounts for 47.55% of single-page and 99.17% of multi-page sessions; 2-Search Engine accounts for 14.63% of single-page versus 0.01% of multi-page; 3-Text-Based Browser accounts for 37.82% of single-page versus 0.82% of multi-page.
  • Bounce rate driven by crawlers and no-referrer traffic: High bounce rates for text-based browsers and search engines are attributed to search engine indexing robots and automated site review tools, which do not support cookies and are therefore recorded as single-page sessions.
  • No-referrer traffic is prominent: Referer_Type "6-No Referrer" accounts for 69.24% of single-page and 34.29% of multi-page sessions, suggesting direct URL entry or bookmark use.
  • Local and internal usage dominates: User_Language_TR = Turkish accounts for 98.47% of single-page and 97.53% of multi-page sessions; User_Location internal (SAU) accounts for 77.56% of single-page and 79.51% of multi-page sessions.
  • Service use and exits (abstract-level result): 40% of users accessed various services, while 50% opted for secure exits, especially when dealing with personal or sensitive information.
  • Association rules: Association rule mining revealed key patterns in frequently accessed services, offering insights into service integration and user preferences.
  • Claimed advantage: The abstract states these findings reveal CAWAL's superiority over conventional methods in precision and efficiency. The paper does not report a head-to-head numerical benchmark against conventional WUM methods in the provided content.

Methodology in Plain English

The researchers implemented the CAWAL model and framework inside CAWIS (Campus Automation Web Information System), Sakarya University's large-scale institutional web portal, which runs on a load-balanced web farm and uses separate subdomains for each service. Unlike traditional WUM, which starts from raw server logs and spends most of its effort on cleaning and session identification, CAWAL collects application-level logs combined with web analytics data, so sessions and page navigation are already accurate and well-structured.

Two augmentation steps enrich the data before mining:

  1. Session augmentation — integrating connected tables, browser data, and user data for a fuller view of user sessions and activities.
  2. User activity augmentation — enriching session and pageview records with login/logout information, server identification, and page duration.

Complex SQL queries were run against one-month data marts (8.5 GB total) to build enriched session and pageview views, exported to CSV. The one-month session file (va_sess5, covering 2022-11-01 to 2022-11-30) contains 1,220,916 records at 235.00 MB; the one-week pageview file (va_page4, covering 2022-11-21 to 2022-11-27) contains 3,158,694 records at 266.00 MB. The session CSV schema includes fields such as Log_ID, Session_ID, Log_Date_Time, User_ID, Session_Login_Status, Logins_During_Period, User_Type, Sex, Age, Age_Group, User_Language_TR, User_Location, Browser_Type, Referer_Type, Landing_Srv_ID, Exit_Srv_ID, Exit_Type, Total_Session_Duration, Avg_Page_Duration, Total_Page_Load, Avg_Page_Load, Page_Count, Visitor_PageView, User_PageView, Service_Count, Page_per_Service, and Visited_Service_IDs. Service-level fields use prefixes: "s_" flags whether a service was visited (1 or 0), "p_" records pages navigated within a service, and "r_" records the ratio of pages in that service to total pages visited.

Four UX-driven analyses were then performed in Python: exit frequency and client attributes, exit methods, transitions between portal services, and association rule mining on user interactions. The bounce-rate analysis used the chi-squared test with the statistic χ² = Σ (Oᵢ − Eᵢ)² / Eᵢ and degrees of freedom (r − 1)(c − 1), with p < 0.05 taken as significant.

On privacy, no synthetic data was used. All data came from the now-retired CAWIS portal via CAWAL, collected under an Internet Services Usage Policy Agreement approved by users and compliant with university regulations and national law, with permission obtained from Sakarya University. Anonymization included irreversible masking of user identities and shifting of timestamps to prevent re-identification.

Why This Matters

Impact on research: The paper shifts attention in web usage mining from improving the pre-processing of messy server logs to improving the collection stage itself. By arguing that structured, application-level enrichment can remove pre-processing altogether, it reframes a long-standing bottleneck (pre-processing consuming substantial WUM time and effort, with results often imprecise) as an architectural rather than an algorithmic problem. It also targets a gap the authors identify: most existing work focuses on static datasets and pays insufficient attention to real-time analysis of streaming data.

Real-world applications:

  • University and institutional portals: CAWIS is a campus automation system with separate subdomains per service, and the findings about multi-page sessions, bounce sources, and secure exits map directly onto improvements for student and staff portals.
  • Enterprise multi-service web portals: The methodology targets load-balanced, multi-server environments where users move between fragmented services, which is where traditional log-based tracking struggles most.
  • Search engine optimization and bot management: The finding that 14.63% of single-page sessions come from search engines and 37.82% from text-based browsers points to concrete bot-traffic and SEO diagnostics.
  • Privacy-sensitive public sector systems: The combination of data ownership, in-house hosting, anonymization by irreversible masking, and timestamp shifting offers a model for organizations that cannot send user data to third-party or cloud analytics.

Industry relevance: The paper frames CAWAL as addressing cost, data security, privacy, and performance concerns the authors associate with big-data and cloud-based WUM solutions, including data breach risk and processing delays. For organizations with high-volume, multi-service web estates, the pitch is a more accurate and cost-effective alternative to conventional log pipelines.

Future Directions

  • Real-time and streaming analysis: The literature review notes most studies still focus on static datasets; extending AWUM to streaming data and its UX implications is an explicitly identified gap.
  • Machine learning and behavior prediction: The paper states that ML and AI techniques in WUM–UX integration are underdeveloped and that CAWAL's enriched datasets can improve prediction and classification models — a direction left open by this study.
  • Privacy-preserving and secure processing methods: The review calls for new methods to process and analyze user data securely as privacy and security concerns become more prominent.
  • Deeper in-page behavior analysis: The paper identifies a gap in analyzing dynamic in-page behaviors and incorporating such data into UX optimization, which the current session- and pageview-level data does not fully cover.

Target Audience

This paper is most useful to web analytics engineers and data scientists building WUM pipelines for large portals; UX researchers and product owners who want behavioral evidence for design decisions; administrators of multi-service institutional platforms (universities, government portals, enterprise intranets) interested in data ownership and privacy-compliant analytics; and HCI researchers studying the intersection of web usage mining and user experience optimization.

Authors’ abstract

Understanding user behavior on the web is increasingly critical for optimizing user experience (UX). This study introduces Augmented Web Usage Mining (AWUM), a methodology designed to enhance web usage mining and improve UX by enriching the interaction data provided by CAWAL (Combined Application Log and Web Analytics), a framework for advanced web analytics. Over 1.2 million session records collected in one month (~8.5GB of data) were processed and transformed into enriched datasets. AWUM analyzes session structures, page requests, service interactions, and exit methods. Results show that 87.16% of sessions involved multiple pages, contributing 98.05% of total pageviews; 40% of users accessed various services and 50% opted for secure exits. Association rule mining revealed patterns of frequently accessed services, highlighting CAWAL's precision and efficiency over conventional methods. AWUM offers a comprehensive understanding of user behavior and strong potential for large-scale UX optimization.

Read the original paper