Skip to content

Latest commit

 

History

12 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

SEC Survey Analysis

Python BERT XGBoost Streamlit

A longitudinal analysis of Texas A&M's Student Engineering Council recruiting surveys, spanning Fall 2021 through Fall 2023, combining NLP sentiment analysis on open-ended student and recruiter feedback with predictive modeling on the structured survey responses.

Try the live demo →

Results at a glance

Best model Random Forest on BERT sentence embeddings + structured features, binary framing
Held-out performance 76.6% accuracy, 0.784 ROC-AUC
Harder framing tested honestly Exact 1-5 rating tops out at 48.6% accuracy, reported as the weaker result rather than dropped
Populations covered Students (4 semesters) and recruiters (5 semesters), analyzed and cross-compared
Abandoned approaches, kept and documented A 9%-accuracy sentiment-only attempt, a 35.4%-accuracy VADER attempt

Three modeling attempts happened here, not one. The first two (9% and 35.4% accuracy) are still in the notebooks rather than deleted, since a model's version history is part of an honest writeup, not something to clean up after the fact.

Objective

Two questions: what do the open-ended comments actually say across three years of career fairs, and can the structured survey responses (attendance, ratings of individual event features) predict a respondent's overall rating of the event.

Data

Four semesters of student feedback (data/StudentData/) and five semesters of recruiter feedback (data/RecruiterData/), each an Excel export with its own column names and, in Fall 2021's case, its own rating scale (1 to 10 instead of 1 to 5). data/cleaned/ holds an intermediate sentiment output from an earlier pass over the Fall 2021 data.

Approach

code/sec-sentiment-analysis-preprocessing.ipynb loads and harmonizes the four student survey files, computes sentiment scores on the free-text feedback columns with a BERT sentiment model (nlptown/bert-base-multilingual-uncased-sentiment), and does the exploratory groundwork (nonresponse testing, column alignment across inconsistent survey instruments). An earlier pass at this notebook tried VADER instead and a rating-scale-only Random Forest trained on Fall 2021 alone (9% accuracy); both are still in the notebook as the record of what didn't work, see reference/CHALLENGES.md.

code/new-sec-sentiment-analysis-work.ipynb is the modeling notebook: BERT sentence embeddings (all-MiniLM-L6-v2) on the aggregated feedback text, reduced with PCA, combined with the structured features, and run through Random Forest and XGBoost in three framings:

Framing Result
Predict exact rating (1-5), regression R² = 0.08 (Random Forest), R² = 0.03 (XGBoost)
Predict exact rating (1-5), classification 48.6% accuracy (XGBoost)
Predict positive vs. not (rating >= 4), classification 76.6% accuracy, 0.784 ROC-AUC (Random Forest)

The binary framing is where the signal actually holds up. Predicting the exact rating turned out to be a harder problem than it looks, largely because 44.8% of all responses across every semester are a top rating, see reference/CHALLENGES.md for why. The same notebook also has an earlier, abandoned attempt at a VADER-plus-text-length Random Forest (35.4% accuracy) that predates the BERT embedding approach and is kept for the same reason.

code/recruiter_analysis.ipynb is the newest addition and looks at the recruiter side, which none of the modeling above touches. Structured recruiter ratings (communication, documentation, virtual platform) pooled across four semesters predict overall recruiter satisfaction at 0.73-0.74 ROC-AUC against a 0.50 baseline, and the effect is almost entirely carried by the communication rating alone. Fall 2023's recruiter survey asks about shuttle waits, check-in, and lunch logistics instead, with no communication question at all, so it can't be pooled with the other four, and gets analyzed separately rather than excluded outright: signage/navigation (r = 0.65, p < 0.0001) and check-in smoothness (r = 0.50, p = 0.0002) together explain 47% of the variance in overall rating at n = 51, while food quality barely matters (r = 0.30). Across all five semesters, worded completely differently every time, the same theme holds: operational clarity drives recruiter satisfaction, amenities don't.

Findings

  • A real decline in company showcase attendance from 2021 to 2022, continuing more modestly into 2023 (reference/outcomes.txt)
  • Nonresponse on the Fall 2022 communication-sentiment question was not random (p = 0.04), the one semester out of four where that held
  • figures/rating_distribution_by_semester.png charts the overall-rating distribution across all four semesters on a common 1-5 scale (Fall 2021 rescaled from its native 1-10)
  • Recruiter satisfaction tracks communication rating almost one-to-one (r = 0.55, n = 238 pooled across four semesters); students show the same relationship but far less consistently (r = 0.19 in Fall 2022, r = 0.58 in Spring 2023), suggesting student satisfaction depends on more than communication quality alone, likely which companies attended and whether relevant roles were available. See code/recruiter_analysis.ipynb for the full comparison.
  • Fall 2023's differently-worded recruiter survey (operations, not communication) shows the same underlying pattern anyway: signage and check-in smoothness explain 47% of overall-rating variance (n = 51), while lunch logistics and food quality don't move the needle much

reference/CHALLENGES.md and reference/ASSUMPTIONS.md cover the survey-instrument inconsistencies, class imbalance, and modeling choices behind these numbers in more detail.

Try it

app/streamlit_app.py is a small Streamlit app built on the binary positive/not-positive framing: input attendance and a handful of structured ratings, and it returns a predicted probability of a positive rating. It's a leaner, structured-features-only version of the notebook model (no BERT text embeddings, to keep the app light and avoid a large model download), so its honest ROC-AUC is 0.70, a bit below the full model's 0.784, stated plainly in the app itself rather than reporting the higher number next to a demo that doesn't achieve it. See app/README.md for how to run it locally or deploy it to Streamlit Cloud.

Running it

pip install pandas numpy scikit-learn xgboost sentence-transformers nltk seaborn matplotlib openpyxl
jupyter notebook code/sec-sentiment-analysis-preprocessing.ipynb

Status

Built independently, analyzing SEC's own recruiting survey data across three academic years.

About

An analysis of the Student Engineering Council's surveys for both students and recruiters across multiple years

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages