In this project, we utilized the AIHub "Capital Region Domestic Travel Log Dataset" to explore travel behavior patterns and develop a prototype system for recommending destinations to foreign tourists in Korea. The dataset comprises approximately 4,000 structured travel logs for individuals visiting the capital region (Seoul, Incheon, and Gyeonggi), including demographic details, GPS-based movement data, activity logs, consumption records, and photos. Our analysis focused on generating actionable insights through Association Rule Mining (ARM), with the objective of uncovering sequential travel behaviors and enabling intelligent recommendation models.
이 프로젝트에서는 AIHub의 "수도권 한국 내 여행로그 데이터셋"을 비롯하여 여러 데이터셋을 통해서 외국인 관광객의 여행 행동 패턴을 분석하고, 여행지 추천을 위한 프로토타입 시스템을 개발하였습니다. 해당 데이터셋은 서울, 인천, 경기 지역을 방문한 여행자의 약 4,000개의 구조화된 여행 로그로 구성되어 있으며, 인구통계 정보, GPS 기반 이동 경로, 활동 기록, 소비 내역, 사진 등의 정보를 포함합니다.
분석의 주요 목적은 연관 규칙 기반 분석(Association Rule Mining, ARM)을 통해 여행지 간의 순차적 이동 패턴을 도출하고, 이를 바탕으로 지능형 추천 모델을 구현하는 것이었습니다.
The primary dataset employed in this study is the “Domestic Travel Log Data (Capital Region)” provided by AIHub, which comprises approximately 4,000 detailed records of travel behavior among visitors to Seoul, Incheon, and Gyeonggi Province. Each record includes multimodal information encompassing demographic attributes, activity logs, GPS trajectories, expenditure details, and photographs. This dataset provides a rich foundation for analyzing sequential travel patterns and user preferences across diverse traveler profiles. AIHub Travel Log Dataset
Demographic Information: The AIHub dataset includes variables such as gender, age group, occupation, income level, and travel party composition. These variables were further enriched using benchmark statistics from the Korea Tourism Organization’s Foreign Visitor Survey, which offers macro-level insights into inbound tourist demographics.
Travel Behavior and Itineraries: The sequential logs of visited POIs were paired with temporal and spatial metadata. To improve interpretability of movement patterns, we referenced data from Seoul Open Data Plaza which provides fine-grained information on public transportation ridership, including subway station usage and transfer patterns.
Environmental and Contextual Variables: To assess the influence of weather on travel behaviors—particularly in shifting preferences between indoor and outdoor
locations—we utilized historical weather observations from the Korea Meteorological Administration (KMA). Variables such as precipitation, temperature, and fine dust concentration were integrated to assess context-aware travel decisions.
Tourist Satisfaction and Intent: Tourist satisfaction levels, preferred activities, and spending patterns were evaluated using annual reports from KTO and cross-validated through behavioral proxies found in the AIHub dataset. These variables are essential for modeling user intent and fine-tuning the recommendation logic.
To uncover underlying patterns in tourist behavior, we applied Association Rule Mining (ARM) using a multi-step process grounded in market basket analysis logic.
The first step involved redefining each travel log as a transaction, analogous to a shopping basket. In our case, the items in the basket were not groceries, but Points of Interest(POIs) that a tourist visited during their trip. For instance, one transaction might look like:
['Incheon Airport', 'Hongdae', 'Namsan Tower']. Each of these POIs is treated as a discrete item, allowing us to analyze co-occurrence relationships across the dataset.
Next, we transformed the transaction list into a binary (one-hot encoded) matrix, where rows represent individual users (or travel logs), and columns represent POIs. A value of 1 indicates the user visited that POI, while 0 indicates they did not. This matrix served as the input for mining frequent patterns. We then applied the Apriori algorithm, a classical ARM technique that efficiently finds frequent itemset using the principle of anti-monotonicity. This principle states that if a set of items is infrequent, any larger set containing it will also be infrequent. With a minimum support threshold set at 25%, the algorithm extracted frequently co-visited POI combinations such as pairs ({Bukchon, Gyeongbokgung}) or triplets. Finally, we generated association rules from the frequent itemsets in the form of logical implications:
“If POI A is visited, POI B is also likely to be visited.”
Each rule was evaluated with three core metrics:
Support: the proportion of logs that contain both POI A and B.
Confidence: the conditional probability of visiting B given that A was visited.
Lift: the ratio of the observed co-occurrence to that expected by chance. A lift greater than 1 indicates a meaningful association. This step-by-step ARM process allowed us to uncover statistically strong and behaviorally meaningful travel patterns, which were later used to support personalized recommendation
systems.
1단계에서는 여행 로그를 거래(transaction)로 재정의했습니다. 각 사용자가 방문한 여
행지를 하나의 '장바구니'로 간주하고, 장소(POI)를 '상품'처럼 취급하여, ['인천공항', '홍대', '남산타워']와 같은 리스트로 구성했습니다.
2단계에서는 이 데이터를 이진 행렬(One-hot Encoding)로 변환했습니다. 행은 사용자,
열은 장소이며, 특정 장소를 방문했는지 여부를 0 또는 1로 표시했습니다.
3단계에서는 Apriori 알고리즘을 통해 자주 함께 등장하는 장소 조합(빈발 항목 집합,
Frequent Itemsets)을 추출했습니다. 최소 지지도(Support)는 25%로 설정하였으며, 이를 기반으로 {북촌, 경복궁}과 같은 장소 쌍이 자주 같이 방문되는지를 파악했습니다.
4단계에서는 도출된 항목 집합으로부터 연관 규칙(Association Rules)을 생성하였습니다.


This simple simulation illustrates how rule-based recommendations can be mapped to user profiles for practical applications in smart tourism apps, kiosks, and travel planning platforms.