This project analyzes data from a few input sources.
Project description
Public data food analysis
Install
pip install public_data_food_analysis_3
Example for data analysis on the Columbia study
import public_data_food_analysis_3.columbia as pdfac
import pandas as pd
Take a brief look on the food logging dataset and the reference information sheet
pdfac.read_logging_data('data/col_test_data').head(2)
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| Unnamed: 0 | original_logtime | desc_text | food_type | PID | |
|---|---|---|---|---|---|
| 0 | 0 | 2021-05-12 02:30:00 +0000 | milk | b | yrt1999 |
| 1 | 1 | 2021-05-12 02:45:00 +0000 | some medication | m | yrt1999 |
pd.read_excel('data/col_test_data/toy_data_17May2021.xlsx').head(2)
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| mCC_ID | Participant_Study_ID | Study Phase | Intervention group (TRE or HABIT) | Start_Day | End_day | Eating_Window_Start | Eating_Window_End | |
|---|---|---|---|---|---|---|---|---|
| 0 | yrt1999 | 2 | S-REM | TRE | 2021-05-12 | 2021-05-14 | 00:00:00 | 23:59:00 |
| 1 | yrt1999 | 2 | T3-INT | TRE | 2021-05-15 | 2021-05-18 | 08:00:00 | 18:00:00 |
make the table that contains extra analytic information that we want
df = pdfac.make_table(pdfac.read_logging_data('data/col_test_data')\
, pd.read_excel('data/col_test_data/toy_data_17May2021.xlsx'))
Participant yrt1999 didn't log any food items in the following day(s):
2021-05-18
Participant yrt2000 didn't log any food items in the following day(s):
2021-05-12
2021-05-13
2021-05-14
df
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| mCC_ID | Participant_Study_ID | Study Phase | Intervention group (TRE or HABIT) | Start_Day | End_day | Eating_Window_Start | Eating_Window_End | phase_duration | caloric_entries | ... | logging_day_counts | %_logging_day_counts | good_logging_days | %_good_logging_days | good_window_days | %_good_window_days | outside_window_days | %_outside_window_days | adherent_days | %_adherent_days | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | yrt1999 | 2 | S-REM | TRE | 2021-05-12 | 2021-05-14 | 00:00:00 | 23:59:00 | 3 days | 7 | ... | 3 | 1.00 | 2.0 | 0.666667 | 3 | 1.00 | 0 | 0.0 | 2 | 0.666667 |
| 1 | yrt1999 | 2 | T3-INT | TRE | 2021-05-15 | 2021-05-18 | 08:00:00 | 18:00:00 | 4 days | 8 | ... | 3 | 0.75 | 2.0 | 0.500000 | 1 | 0.25 | 2 | 0.5 | 1 | 0.250000 |
| 2 | yrt2000 | 3 | T3-INT | TRE | 2021-05-12 | 2021-05-14 | 08:00:00 | 16:00:00 | 3 days | 0 | ... | 0 | 0.00 | 0.0 | 0.000000 | 0 | 0.00 | 0 | 0.0 | 0 | 0.000000 |
3 rows × 26 columns
df.iloc[0]
mCC_ID yrt1999
Participant_Study_ID 2
Study Phase S-REM
Intervention group (TRE or HABIT) TRE
Start_Day 2021-05-12 00:00:00
End_day 2021-05-14 00:00:00
Eating_Window_Start 00:00:00
Eating_Window_End 23:59:00
phase_duration 3 days 00:00:00
caloric_entries 7
mean_daily_eating_window 13.75
std_daily_eating_window 11.986972
earliest_entry 4.5
2.5% 4.5375
97.5% 27.5625
duration mid 95% 23.025
logging_day_counts 3
%_logging_day_counts 1.0
good_logging_days 2.0
%_good_logging_days 0.666667
good_window_days 3
%_good_window_days 1.0
outside_window_days 0
%_outside_window_days 0.0
adherent_days 2
%_adherent_days 0.666667
Name: 0, dtype: object
df.iloc[1]
mCC_ID yrt1999
Participant_Study_ID 2
Study Phase T3-INT
Intervention group (TRE or HABIT) TRE
Start_Day 2021-05-15 00:00:00
End_day 2021-05-18 00:00:00
Eating_Window_Start 08:00:00
Eating_Window_End 18:00:00
phase_duration 4 days 00:00:00
caloric_entries 8
mean_daily_eating_window 8.666667
std_daily_eating_window 8.504901
earliest_entry 7.5
2.5% 7.7
97.5% 23.9
duration mid 95% 16.2
logging_day_counts 3
%_logging_day_counts 0.75
good_logging_days 2.0
%_good_logging_days 0.5
good_window_days 1
%_good_window_days 0.25
outside_window_days 2
%_outside_window_days 0.5
adherent_days 1
%_adherent_days 0.25
Name: 1, dtype: object
Example for data analysis using public data food analysis module
import public_data_food_analysis_3.core as pdfaco
import pandas as pd
take a look at the original dataset
df = pdfaco.universal_key('data/test_food_details.csv')
df.head(2)
read the csv file successfully.
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| Unnamed: 0 | ID | unique_code | research_info_id | desc_text | food_type | original_logtime | foodimage_file_name | |
|---|---|---|---|---|---|---|---|---|
| 0 | 1340147 | 7572733 | alqt14018795225 | 150 | Water | w | 2017-12-08 17:30:00+00:00 | NaN |
| 1 | 1340148 | 411111 | alqt14018795225 | 150 | Coffee White | b | 2017-12-09 00:01:00+00:00 | NaN |
preprocess the data to have extra basic features
df = pdfaco.load_public_data(df)
df.head(2)
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| Unnamed: 0 | ID | unique_code | research_info_id | desc_text | food_type | original_logtime | original_logtime_notz | date | local_time | time | week_from_start | year | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | 1340147 | 7572733 | alqt14018795225 | 150 | Water | w | 2017-12-08 17:30:00+00:00 | 2017-12-08 17:30:00+00:00 | 2017-12-08 | 17.500000 | 17:30:00 | 1 | 2017 |
| 1 | 1340148 | 411111 | alqt14018795225 | 150 | Coffee White | b | 2017-12-09 00:01:00+00:00 | 2017-12-09 00:01:00+00:00 | 2017-12-08 | 24.016667 | 00:01:00 | 1 | 2017 |
do a brief annalysis
df = pdfaco.summarize_data(df, 'local_time', 'unique_code')
df.head(2)
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
| unique_code | num_days | num_total_items | num_f_n_b | num_medications | num_water | breakfast_avg | breakfast_std | dinner_avg | dinner_std | eating_win_avg | eating_win_std | good_logging_count | breakfast variation (90%-10%) | dinner variation (90%-10%) | 2.5% | 95% | duration mid 95% | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 0 | alqt1148284857 | 13 | 149 | 96 | 19 | 34 | 7.821795 | 6.710717 | 23.485897 | 4.869082 | 15.664103 | -1.841635 | 146 | 2.966667 | 9.666667 | 4.535000 | 26.813333 | 22.636667 |
| 1 | alqt14018795225 | 64 | 488 | 484 | 3 | 1 | 7.525781 | 5.434563 | 25.858594 | 3.374839 | 18.332813 | -2.059723 | 484 | 13.450000 | 3.100000 | 4.183333 | 27.438333 | 23.416667 |
df.iloc[0]
unique_code alqt1148284857
num_days 13
num_total_items 149
num_f_n_b 96
num_medications 19
num_water 34
breakfast_avg 7.821795
breakfast_std 6.710717
dinner_avg 23.485897
dinner_std 4.869082
eating_win_avg 15.664103
eating_win_std -1.841635
good_logging_count 146
breakfast variation (90%-10%) 2.966667
dinner variation (90%-10%) 9.666667
2.5% 4.535
95% 26.813333
duration mid 95% 22.636667
Name: 0, dtype: object
df.iloc[1]
unique_code alqt14018795225
num_days 64
num_total_items 488
num_f_n_b 484
num_medications 3
num_water 1
breakfast_avg 7.525781
breakfast_std 5.434563
dinner_avg 25.858594
dinner_std 3.374839
eating_win_avg 18.332813
eating_win_std -2.059723
good_logging_count 484
breakfast variation (90%-10%) 13.45
dinner variation (90%-10%) 3.1
2.5% 4.183333
95% 27.438333
duration mid 95% 23.416667
Name: 1, dtype: object
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file public_data_food_analysis_3-0.2.7.tar.gz.
File metadata
- Download URL: public_data_food_analysis_3-0.2.7.tar.gz
- Upload date:
- Size: 112.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/3.7.1 importlib_metadata/4.11.2 pkginfo/1.8.2 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.59.0 CPython/3.8.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b0c5fe135b2d3478781948f2d7a7ed43cb11f2a090170a39024135f8d82a6167
|
|
| MD5 |
04ced0ccd8865fdf084779a67cac24b3
|
|
| BLAKE2b-256 |
0ff720d190fc76725d605f19cafbff3f58df160bd990b77afb396f6ad5622d79
|
File details
Details for the file public_data_food_analysis_3-0.2.7-py3-none-any.whl.
File metadata
- Download URL: public_data_food_analysis_3-0.2.7-py3-none-any.whl
- Upload date:
- Size: 109.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/3.7.1 importlib_metadata/4.11.2 pkginfo/1.8.2 requests/2.25.1 requests-toolbelt/0.9.1 tqdm/4.59.0 CPython/3.8.8
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
221fc212ef9c04d39086c0c308264316a05856326af4e332f71501e83b3d60a0
|
|
| MD5 |
d3243b0655da6eb6cb10f284e98b77f5
|
|
| BLAKE2b-256 |
97f012e5cd99301f2cc8fe594ce434110b7ea90d6cf1cc4c01fd9212922d9d18
|