
https://nbviewer.jupyter.org/gist/oguzhanilter/cf65dde581b67e7a6ea538f6eeba8759
Project Description
In this project, we want to analyse the users of the Los Angeles Metro Bike-Sharing Service and their usage preferences depending on time and locations of the city in order to clarify:
- Prefered time intervals
- Effect of crime rate in the area
- Traffic factor
- General user statistics
At the end of the project we aim to determine some alternative locations for Bike-Stations based on users’ route preferences.
Used Datasets
We obtained the datasets from Kaggle along with Metro Bikeshare Service official website (https://bikeshare.metro.net/about/data/).
The Crime Dataset contains crime type, place and time in the city. By using the Crime Dataset, we want to obtain general safety profiles of the locations.
Crimes in Los Angeles Dataset has 1584316 rows, and 26 columns.
Information on every column and their types:
DR Number int64
Date Reported object
Date Occurred object
Time Occurred int64
Area ID int64
Area Name object
Reporting District int64
Crime Code int64
Crime Code Description object
MO Codes object
Victim Age float64
Victim Sex object
Victim Descent object
Premise Code float64
Premise Description object
Weapon Used Code float64
Weapon Description object
Status Code object
Status Description object
Crime Code 1 float64
Crime Code 2 float64
Crime Code 3 float64
Crime Code 4 float64
Address object
Cross Street object
Location object
dtype: object
The Traffic Collision Dataset contains information about traffic accidents such as victim, time, location etc. This is the biggest dataset that we use but most of the accidents are not related to our topics. Therefore this dataset will be downsized by keeping only bike-related accidents.
Traffic Accidents in Los Angeles Dataset has 463819 rows, and 18 columns.
Information on every column and their types:
DR Number int64
Date Reported object
Date Occurred object
Time Occurred int64
Area ID int64
Area Name object
Reporting District int64
Crime Code int64
Crime Code Description object
MO Codes object
Victim Age float64
Victim Sex object
Victim Descent object
Premise Code float64
Premise Description object
Address object
Cross Street object
Location object
dtype: object
Metro Bike Sharing Service Dataset is our main dataset and contains information about every trip defined with the ancillary information about time, location, type of costumer and Bike-Station. Locations of Bike-Stations have been provided both in coordinates and ID of station. A supplementary document contains information about those IDs.
Bike Share Dataset has 132427 rows, and 16 columns.
Information on every column and their types:
Trip ID int64
Duration int64
Start Time object
End Time object
Starting Station ID float64
Starting Station Latitude float64
Starting Station Longitude float64
Ending Station ID float64
Ending Station Latitude float64
Ending Station Longitude float64
Bike ID float64
Plan Duration float64
Trip Route Category object
Passholder Type object
Starting Lat-Long object
Ending Lat-Long object
dtype: object
One problem about the main dataset is, although it contains Los Angeles in the name, the bike-sharing system is applicable only in a small area. Therefore it becomes harder to combine with different datasets as they cover more area and lack detail in area based information.
Center of Crimes and Center of Accidents
City of Los Angeles consists of 21 sub-areas each controlled by different police stations. When the fusion of two datasets (The Traffic Collision in Los Angeles Dataset and The Crime in Los Angeles Dataset) is considered, the number of locations a crime or accident can happen is very big. This makes it very hard to keep track of the effects of every individual crime or accident on the Bike-Sharing Service users’ preferences. Therefore we will create a Center of Crimes and a Center of Accidents for every sub-area.
Substep 1: Pre-Processing
The Traffic Collision in Los Angeles and
The Crime in Los Angeles Datasets.
We will only use a segment of the Traffic Collision Dataset. If we examine the MO Code explanations, there are bike related accidents. We will use only the rows which contain these accidents. The MO Codes are: 0345, 3008, 1223, 3016, 3017, 3018, 3021
The longitude and latitude values were strings in both The Traffic Collision in Los Angeles Dataset and
The Crime in Los Angeles Dataset. So we needed to convert them into integers.
In addition to locations of the Centers, we will need the total number of crime and accidents in the 21 sub-areas. Calculation of Centers is as follows:

We also divided the Time of Day into 6 sections. This way we wont need to deal with continious time series.
Crime Ratio for Time of Day:
Noon 0.343702
Evening 0.281171
Late Night 0.171482
Rush Hour Afternoon 0.106698
Rush Hour Morning 0.067934
Early Birds 0.029013
Name: Time of Day
dtype: float64
Metro Bike Sharing Service Dataset
The time of the day representions does not match between datasets, therefore we will create new columns in order to represent the time in the same format as in The Traffic Collision in Los Angeles Dataset and The Crime in Los Angeles Dataset. In these datasets time is just an integer, which is created in following manner: HHMM. For example 14:02 (2:02 PM) is reprensted as 1402.
In addition to time representation, we will classify the ‘Start Time’ into 6 time of day categories same as the other datasets.
Also we will create a column which represents the day of the week for every trip. However for the sake of simplicity we will classify this information into two categories: Weekday and Weekend.
Weekday 96594
Weekend 35833
In addition to that, we want to see the locations of the Stations but we can not find that information by using Latitude and Longitude because it does not give enough accuracy.However, we can still use those coordinates to find the distance.
extended_bike[‘distance’] = [geodesic((row[‘Starting Station Latitude’], row[‘Starting Station Longitude’]),(row[‘Ending Station Latitude’], row[‘Ending Station Longitude’])).miles*1.60934 for index,row in extended_bike.iterrows()]
Final Shapes of Datasets
Metro Bike Sharing Service Dataset
New Bike Share Dataset has 131276 rows, and 25 columns
Information on every column and their types:
Trip ID int64
Duration int64
Start Time object
End Time object
Starting Station ID float64
Starting Station Latitude float64
Starting Station Longitude float64
Ending Station ID float64
Ending Station Latitude float64
Ending Station Longitude float64
Bike ID float64
Plan Duration float64
Trip Route Category object
Passholder Type object
Starting Lat-Long object
Ending Lat-Long object
dtype: object
The Traffic Collision in Los Angeles Dataset
New Traffic Dataset has 15976 rows, and 21 columns
Information on every column and their types:
Trip ID int64
Duration int64
Start Time object
End Time object
Starting Station ID float64
Starting Station Latitude float64
Starting Station Longitude float64
Ending Station ID float64
Ending Station Latitude float64
Ending Station Longitude float64
Bike ID float64
Plan Duration float64
Trip Route Category object
Passholder Type object
Starting Lat-Long object
Ending Lat-Long object
dtype: object
The Crime in Los Angeles Dataset.
New Crime Dataset has 1578834 rows, and 29 columns
Information on every column and their types:
Trip ID int64
Duration int64
Start Time object
End Time object
Starting Station ID float64
Starting Station Latitude float64
Starting Station Longitude float64
Ending Station ID float64
Ending Station Latitude float64
Ending Station Longitude float64
Bike ID float64
Plan Duration float64
Trip Route Category object
Passholder Type object
Starting Lat-Long object
Ending Lat-Long object
dtype: object
Assumption
We will assume that the rate of crime incidents for the region of a Bike-station will be equal to the number of crimes of nearest Center of Crimes divided by distance between The Center of Crimes and that station. Likewise, the rate of traffic accidents will be equal to the number of accidents of nearest Center of Accidents divided by distance. So that we will have the general profile of every Bike-station region.
Substep 2: Data Exploration
In this section we will visualize our datasets and try to create some hypothesis in order to find some correlations between them
Locations of Bike Stations and Crime Rate Circles :

Distance and Pass-Holder Type:

Variation of Average Duration as Time of Day Changes:

Ratio of Users vs Time of Day:

Rate of Crimes vs Time of Day:

Rate of Accidents vs Time of Day:

Users, Crime, and Accidents per station:



Substep 3: Hypothesis Testing
1. Are the long travel times more likely in Rush Hour?
Null Hypothesis (H0) : Longer travels are more likely in Rush Hour.

P-value is much lower than 0.05. Therefore we concluded that the null hypothesis is not correct. Longer travels are not more likely in rush hour.
2. Is Pass Holder Type a significant indicator for Long Travelling Distance?
Null Hypothesis (H0) : People with monthly pass go shorter distances because they have a daily routine.

P-value is much lower than 0.05. Therefore we concluded that the null hypothesis is not correct. People with montly pass does not necessarily go shorter distances.
3. Is Long Duration more common in High Crime Areas?
Null Hypothesis (H0) : Areas where people ride bikes for longer times have higher crime rates

Although the P-value is around 0.01, we will assume that the Null Hypothesis holds. Because the value of Crime Rate is also an assumption for every area. Normally, we would only accept P-values above 0.05 but int his particular case we will accept 0.01.
4. Is Long Riding Distance a significant factor for Accident Rate?
Null Hypothesis (H0): The chance of accidents increases as the riding distance gets longer.

Result => (statistic=0.10329975155008089, pvalue=0.917725247586795)
As the P-value is way above 0.05, we can confidently conclude that the Null Hypothesis holds. When people ride for further distances, chance of an accident increases.
5. Effect of long duration ride on Accident Rate
Null Hypothesis (H0): The chance of accidents increases as the duration gets longer

Similar with number 4, the P-value is way above 0.05. We can confidently conclude that the Null Hypothesis holds. When people ride for longer durations, chance of an accident increases.
The Relationship Between Crime Rate and Accident Rate – Linear Regression
In this part we are asked to build a Linear regression model in order to predict some information in future. Our values are the Crime Rate and Accident Rate. Also by looking this we can conclude that these two values are linearly correlated.



(-0.6378231281388407, 0.026378162963240314)
The bias of model is -0.6378 and the slope is 0.0263