Research Data & Computing
Data management systems, research data services, and high-performance and cloud computing.
- HPC, Linux & AWS LightSail for Research
- Research data management & IRB support
- Sequence data analysis (SRA / NCBI)
Samah Alshrief, Ph.D.
AI data scientist, research-computing specialist, and higher-education policy analyst.
I help universities and researchers turn complex questions into working data and computing. I build research data services, private AI tools, and the infrastructure that makes computational research possible. A particular focus: connecting small institutions to the national computing resources their researchers already qualify for.
About
I'm a data scientist and research-computing specialist at Seton Hall University, where I help faculty, students, and whole institutions turn stalled research questions into working data pipelines, computing workflows, and evidence they can stand behind. I helped launch the university's first computational research projects, and I now co-lead its Research Computing Initiative.
My work sits where three fields meet. I hold a Ph.D. in Higher Education Leadership, Management, and Policy, with a concentration in research, assessment, and program evaluation, plus a Master of Science in Data Science and Engineering and a Master of Education. That mix is the point: I can build the technical thing and explain, in plain terms, why it matters to the people who fund it.
The through line is evaluation and evidence. I trained in quantitative methods at ICPSR's Summer Program at the University of Michigan, and my policy work runs from program evaluation and budget analysis to interrupted time-series studies of enrollment policy. Along the way I've contributed to more than $136,000 in funded projects, supported IRB applications for research teams, delivered over 100 data consultations, and coordinated 43 concurrent research projects. I serve as Seton Hall's ICPSR Official Representative and a Women in Data Science (WiDS) ambassador, and I mentor people moving into data careers.
Based in New Jersey and New York City. I work with teams on site and remotely.
Data management systems, research data services, and high-performance and cloud computing.
Program evaluation and policy research grounded in rigorous methods.
A practitioner inside the university, building research-computing capacity and training people to use it.
Document-grounded AI built on institutional data, with attention to privacy and ethics.
Work with me
Advising, training, and building for universities, research teams, government agencies, and organizations working with data and AI.
For small and mid-sized institutions without a research computing office: I connect your researchers to national resources like the OSPool and NSF ACCESS (much of it free to use), pick the right venue for each workload, and get the first pilots running so the capability outlasts my visit. I know both sides of this gap because I'm building the same capacity inside a university now.
Research data strategy, research-computing setup (HPC and cloud), data governance, and IRB-aligned data management. Helping teams stand up the infrastructure and practices that make computational research possible.
Hands-on sessions in data science and research tools (R, Python, STATA, ArcGIS, Atlas.ti, Qualtrics, research data management, and AI literacy), tailored to your team's projects and skill level.
Design and deployment of document-grounded, private AI tools over your institution's own data: retrieval-augmented systems that keep sensitive material in-house, as in my syllabus-analysis platform.
Talks and guest sessions on AI in higher education, research computing, private AI, and data ethics, for conferences, campuses, and organizations.
For agencies, institutions, and nonprofits: evaluation design, quantitative and qualitative analysis, and plain-language reporting on what a program actually does. My doctorate concentrated in research, assessment, and program evaluation, with ICPSR-trained methods behind it.
Methods and data support for doctoral students and research teams: study and survey design, quantitative and qualitative analysis, and getting messy data ready for a dissertation or a manuscript.
One-on-one mentoring for students and early-career researchers building data, computing, and AI skills, from a first analysis to a research-computing workflow they can run on their own.
Selected Work
A sample of applied projects across AI, research computing, and data science. The full record is in my CV.
Designed and deployed a private AI app for Business School faculty to analyze syllabi and generate document-grounded feedback, built with Python, Streamlit, ChromaDB, and a retrieval-augmented pipeline.
Part of the team forming a Research Computing Committee and an internal grant proposal for HPC clusters, high-performance storage, and scalable resources to support computational research across disciplines.
Helped launch the university's first cloud research-computing project (RStudio on AWS LightSail for Research) and taught sequence-data analysis using SRA and NCBI toolkits.
Built predictive models on 3D biomechanics data of baseball players, using regression, decision trees, random forests, and SVMs, to identify predictors of speed and reduce injury risk.
Coordinated 43 concurrent research projects, delivering data-management and analysis training across quantitative, qualitative, ArcGIS, and social-media methods, with one-on-one student consultations.
Served as data specialist for the School of Diplomacy, supporting Atlas.ti and Stata training and building lesson content with the policy librarian for international-relations research.
Background
Teaching
University courses and hands-on workshops in data science, research methods, and computing. Click a course to request its syllabus.
Writing & Talks
Selected talks below. Articles and commentary on AI, research computing, and higher education are coming here. Subscribe or reach out to follow along.
Why "we ran the program and things got better" isn't an answer, the "compared to what" problem at the center of every evaluation, and which designs a real budget can actually support.
Read more →Your researchers likely qualify for national computing resources like the OSPool and NSF ACCESS at no cost. Why small campuses miss out, and how to get a first researcher running.
Read more →What consent actually covers, why dropping the name column doesn't de-identify anything, when secondary use needs another look, and what to check before data touches an AI tool.
Read more →The core skills worth teaching faculty and staff, a few durable questions to ask of any AI answer, and how to run a workshop that changes habits instead of just impressing the room.
Read more →What each required section of a DMP actually asks for, the mistakes that get plans flagged, and a reusable outline you can adapt per grant.
Read more →What private, document-grounded AI is, when a university should use it instead of a public chatbot, and how to build one responsibly.
Read more →Questions
A few things people usually want to know before reaching out.
Universities and colleges, research teams, government agencies, nonprofits, and doctoral students. The common thread is work that runs on data, research computing, or AI, whether that's setting up research data services, evaluating a program, or building a private AI tool.
Yes, that's a focus of my work. Many small and mid-sized campuses already qualify for national resources like the OSPool and NSF ACCESS, much of it free to use, but never get set up. I connect your researchers to the right resource, match each workload to the right venue, and get first pilots running so the capability stays after the project ends.
Private AI is a document-grounded system built over your institution's own material, so answers come from your data and sensitive content stays in-house. A public chatbot answers from general training data and sends your prompts to an outside service. When the material is confidential or governed by policy, a private, retrieval-augmented setup is usually the safer choice.
I'm based in New Jersey and New York City and work with teams remotely. Workshops and talks can be delivered on site or online depending on what suits your group.
Yes. Sessions in data tools, research data management, and AI literacy are tailored to your team's projects and skill level, including groups that are new to the material. The aim is habits people keep using, not a one-time demo.
Start with a short discovery call to scope your project or event. From there, engagements usually take the shape of a workshop or talk, a scoped project, or ongoing advisory. Use the contact form to tell me a little about what you need and I'll reply, usually within a couple of business days.
Contact
Tell me a little about your project or event and I'll reply, usually within a couple of business days.